Deterministic and unspecified results

The results of GeoDMS operators are deterministic: one version of GeoDMS, given the same configuration and the same source data, gives the same results every time. Some results are also unspecified: the documentation of the operator names the values a result may take without saying which one it is. This page says what both words mean, and what a configuration or a test must do with an unspecified result.

Deterministic

A result is deterministic when, with one and the same version of GeoDMS, the same configuration and the same source data, it is the same:

  • in every run, in the GUI or in GeoDmsRun;
  • within one run, whatever the order in which items and their tiles are requested, calculated, calculated side by side in parallel, released from memory and calculated again.

Deterministic says nothing about another version of GeoDMS. Nor does it say anything about other input, and that includes input that seems unrelated: where a choice is unspecified (below), adding or removing objects elsewhere can change which value an operator chooses, because it changes how the operator organises its search.

Specified and unspecified

A result is specified when the page of the operator says exactly which value it is. For instance, Point_in_ranked_polygon gives, of the polygons that contain a point, the one with the highest rank, and of polygons of equal rank the one with the lowest index. A specified result changes between versions only as a documented correction, marked on the page with “Since GeoDMS x.y.z”.

A result is unspecified where the page names the values it may take but not which one of them it is. For instance, for a point in several overlapping polygons, Point_in_polygon gives one of those polygons. An unspecified result is still deterministic. But which of the allowed values it is:

  • can differ between versions of GeoDMS. In 20.22.1 it did, for a point in several identical polygons, when the spatial index that point_in_polygon uses started to split its nodes at another size (#1289);
  • can differ within one version when other parts of the input change (see above).

So a configuration, a test, or a comparison between the results of two versions must accept each of the allowed values. Where the choice matters, make it explicit: use Point_in_ranked_polygon instead of Point_in_polygon, make the candidates differ, select them before the operator, or use the filter of Connect_eq or Connect_ne.

A choice is left unspecified where a rule would cost time for every user of the operator. point_in_polygon stops at the first polygon it finds that contains the point; to give the polygon with the lowest index, it would have to look at all the others as well, which made it about 7 times slower on polygons that do not overlap. It also leaves room to make a search faster in a later version.

Unspecified results

operator unspecified specified
Point_in_polygon which of several polygons that contain a point the polygon, for a point in one polygon only; null for a point in none
Connect, Connect_eq, Connect_ne, Capacitated_connect which of several equally near points, arcs or polygons is connected the distance of the connection
Connect_info, Connect_info_eq, Connect_info_ne which of several equally near arcs (or points) is taken, and so ArcID or point_rel, CutPoint, InArc, InSegm and SegmID dist
Connect_neighbour which of several equally near points is taken the distance to it

dist_info gives only the distance, which is the same for each of equally near candidates, so its result is specified.

See also