Deepening III, Python and GeoDMS side by side

learning objective: reading the same aggregation in Python and in the GeoDMS, mapping the Python idioms onto GeoDMS operators, and knowing what makes each implementation fast or slow

introduction

Module 2D, GeoDMS and Python shows how the GeoDMS and Python exchange data. This deepening looks at the two from another angle: one calculation, written once in Python and once in the GeoDMS, run on real model output of 173 million rows. You will see how the idioms of one language translate into the other, and what the difference in speed is made of.

The case comes from NetworkModel_PBL, an accessibility model made for the Netherlands Environmental Assessment Agency (PBL). For every origin-destination pair (OD pair) it calculates the Pareto-optimal travel options: options for which no alternative is better on all criteria, such as travel time and price. An option is a public transport chain with pre- and post-transport, or a direct trip on foot, by bike or by car. The output step writes one CSV file per block of origins (450 blocks) and per departure moment (07:00, 07:15, 07:30 and 07:45), with one row per option. For the morning of 6 October 2026 that is 1,800 files, 23.3 GB, and 173,000,050 rows.

Both implementations are in use. The GeoDMS code is part of the output step and makes the counts in every run; the Python script makes the same counts for output that was written before the GeoDMS code existed.

This page assumes you know relations and aggregations (Module 1c, Learning the basic concepts of GeoDMS, calculations over multiple domains), for_each and indirect expressions (Module 3, Meta scripting ‐ templates and for_each), and the performance concepts of Deepening II, Performance and memory.

the question

Each row of an output file has, among its 31 columns, the origin (OrgName), the destination (DestName), the modes used (ModeUsed_<moment>), and whether the option needs a car or a bike (NeedsCar_<moment>, NeedsBike_<moment>, 0 or 1):

OrgName;DestName;Traveltime_07h00m;Price_07h00m;Price_Augm_07h00m;ModeUsed_07h00m;NeedsCar_07h00m;NeedsBike_07h00m;...
PandBlok_65575;Rheden;40.133335;5.21;8.23;W_BBR_W;0;0;...

ModeUsed is WW, CC or AA for a direct trip on foot, by bike or by car. A public transport chain has underscores, as in W_BBR_W: walk to the first stop, three public transport legs, walk from the last stop.

Two small tables summarise such output at a glance. The model’s users are Dutch, and so are the labels:

  • od_tellingen.csv (counts of OD rows): the number of rows per kind of option, per departure moment and in total. The kinds are alle (all), OV-keten (a public transport chain), WW, CC and AA.
  • unieke_od_tellingen.csv (counts of unique OD pairs): the number of unique OD pairs with at least one option of that kind, or with at least one option for a group of travellers: zonder auto (without a car: NeedsCar is 0) and zonder auto en fiets (without a car and a bike). Per departure moment, and for the four moments together (samen).

For the morning of 6 October 2026:

soort (kind) 07h00m 07h15m 07h30m 07h45m totaal
alle 43,214,954 43,220,795 43,320,686 43,243,615 173,000,050
OV-keten 3,821,320 3,826,076 3,941,686 3,854,368 15,443,450
WW 107,284 107,284 107,284 107,284 429,136
CC 836,827 836,827 836,829 836,828 3,347,311
AA 38,449,523 38,450,608 38,434,887 38,445,135 153,780,153
groep (group) 07h00m 07h15m 07h30m 07h45m samen
alle 14,443,660 14,444,579 14,443,642 14,445,007 14,448,616
OV-keten 1,533,524 1,539,869 1,580,944 1,549,725 2,087,775
WW 107,284 107,284 107,284 107,284 107,284
CC 836,827 836,827 836,829 836,828 836,830
AA 14,333,586 14,333,004 14,331,709 14,331,660 14,387,855
zonder auto 2,040,917 2,040,659 2,078,823 2,048,502 2,516,369
zonder auto en fiets 1,589,215 1,594,529 1,635,229 1,603,859 2,132,806

So 173 million options connect 14.4 million OD pairs, and 2.5 million of those pairs can be travelled without a car at one of the four moments.

The code below keeps the Dutch names of the model. A small glossary: tellingen = counts, rijen = rows, paren = pairs, soort = kind, groep = group, keten = chain, OV = public transport, blok = block, moment = departure moment, uitvoermap = output folder.

One property of the data makes both implementations simple: every origin belongs to exactly one block, so every OD pair occurs in one block only. The counts of the blocks can just be added up. Only the union over the four departure moments needs care, and that union stays within a block.

the Python implementation

The script handles one block at a time. It streams the rows of the four files of the block through csv.reader, counts the rows per kind in a Counter, and collects the pairs per group in Python sets of (origin, destination) tuples. The union over the moments is a set union. Because the blocks are independent, a multiprocessing.Pool spreads them over 8 processes, and the main process adds up the results:

import csv, glob, os, re
from collections import Counter, defaultdict
from multiprocessing import Pool

NAME_RE = re.compile(r'tt_\d{8}_(\d\dh\d\dm)_(\d+)of(\d+)\.csv$')

def blok(files):
	rijen, paren, samen = Counter(), Counter(), defaultdict(set)
	for p in files:
		m = NAME_RE.search(os.path.basename(p)).group(1)
		per = defaultdict(set)
		with open(p, newline='', encoding='utf-8') as fh:
			r = csv.reader(fh, delimiter=';')
			h = next(r)
			im, ic, ib = h.index('ModeUsed_' + m), h.index('NeedsCar_' + m), h.index('NeedsBike_' + m)
			for row in r:
				if not row:
					continue
				od = (row[0], row[1])
				mode = row[im]
				soort = 'OV-keten' if '_' in mode else mode
				rijen[(m, 'alle')] += 1
				rijen[(m, soort)] += 1
				per['alle'].add(od)
				per[soort].add(od)
				if row[ic] == '0':
					per['zonder auto'].add(od)
					if row[ib] == '0':
						per['zonder auto en fiets'].add(od)
		for g, s in per.items():
			paren[(m, g)] += len(s)
			samen[g] |= s
	return rijen, paren, Counter({g: len(s) for g, s in samen.items()})

# in main(): group the files per block, count the blocks in parallel, add up
per_blok = defaultdict(list)
for p in glob.glob(os.path.join(uitvoermap, 'PerBlock', 'tt_*.csv')):
	per_blok[NAME_RE.search(os.path.basename(p)).group(2)].append(p)
R, P, S = Counter(), Counter(), Counter()
with Pool(8) as pool:
	for r, p, s in pool.imap_unordered(blok, per_blok.values()):
		R.update(r); P.update(p); S.update(s)
# ... and write R, P and S as od_tellingen.csv and unieke_od_tellingen.csv

The full script, which also refuses to count when a block or a moment is missing, is batch/OdTellingen.py.

Note how much of this code has nothing to do with the question itself: finding the files by name, finding the columns by header, parsing 31 text columns per row to use 5 of them, and identifying an OD pair by two strings.

the GeoDMS implementation

In the GeoDMS the counts are made in the output step itself, from the final result of each departure moment. That result is in memory anyway, because the CSV file is written from it. The code has three levels: per departure moment, per block, and over all blocks.

Together, the code blocks of this section form one complete configuration. Put the two classifications and the three Tellingen containers in the skeleton below, at the places marked see above and see below, and it runs as it is, on small test data instead of the model’s results (see try it yourself!).

the classifications and the model around the counts

The kinds and the groups are two small classifications. The code uses their order and element names; the files show their labels:

unit<uint8> OD_Tellingen_Soort : nrofrows = 5
{
	attribute<string> name  : ['alle', 'OV_keten', 'WW', 'CC', 'AA'];
	attribute<string> label : ['alle', 'OV-keten', 'WW', 'CC', 'AA'];
}
unit<uint8> OD_Tellingen_Groep : nrofrows = 7
{
	attribute<string> name  : ['alle', 'OV_keten', 'WW', 'CC', 'AA', 'zonder_auto', 'zonder_auto_en_fiets'];
	attribute<string> label : ['alle', 'OV-keten', 'WW', 'CC', 'AA', 'zonder auto', 'zonder auto en fiets'];
}

The counts sit in the container structure of the model. The skeleton below has that structure, and declares the items the counts use, with test data where the model has its own calculations:

container Tellingen_Example
{
	container Classifications
	{
		// unit<uint8> OD_Tellingen_Soort and unit<uint8> OD_Tellingen_Groep: see above
	}
	container ModelParameters
	{
		container Advanced
		{
			// in the model: the departure moments chosen from the sample day
			unit<uint32> MeasureMoments : nrofrows = 2
			{
				attribute<string> Name                : ['At_07h00m00s', 'At_07h15m00s'];
				attribute<string> TemplatableTextShrt : ['07h00m', '07h15m'];
			}
			parameter<string> OutputDir := '%LocalDataProjDir%/Output';
		}
	}
	container NetworkSetup : using = "Classifications;ModelParameters"
	{
		container impl
		{
			// in the model: 450 blocks of origins
			unit<uint32> OrgBlockDomain : nrofrows = 2
			{
				attribute<string> name : ['Block_1of2', 'Block_2of2'];
			}
		}
		container ConfigurationPerBlock := for_each_ne(impl/OrgBlockDomain/name, 'ConfigurationPerOrgBlock_T(' + string(id(impl/OrgBlockDomain)) + ')')
		{
			container Generate_Output
			{
				parameter<string> OUTPUT_Generate_PublicTransport_fullOD_long_CSVFiles := =asList(NetworkSetup/Impl/OrgBlockDomain/name+'/PublicTransport/Generate_Output/Traveltime_ForEachDepTime_Fence/Generate',' + ') + ' + OD_Tellingen/Generate';

				// container OD_Tellingen: over all blocks, see below
			}
		}
		Template ConfigurationPerOrgBlock_T
		{
			parameter<NetworkSetup/impl/OrgBlockDomain> BlockDomain_id;
			///
			container PublicTransport
			{
				// in the model: all stops, origins and destinations, and the modes of the feed plus walking, cycling and car
				unit<uint32> Places : nrofrows = 50;
				// one element per (origin, destination) pair
				unit<uint64> OD_code := combine_unit_uint64(Places, Places);
				unit<uint8>  Modes  : nrofrows = 3
				{
					attribute<string> name : ['Walking', 'Cycling', 'Car'];
					container V := for_each_nedv(name, String(ID(.))+'[..]', void, .);
				}
				container Generate_Output
				{
					container Traveltime_ForEachDepTime
					{
						parameter<string> Generate := ='add('+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Generate',',')+')';

						// container Tellingen: per block, see below
					}
					container Traveltime_ForEachDepTime_Fence := PhaseContainer(Traveltime_ForEachDepTime, 'Results for all departure times in '+NetworkSetup/Impl/OrgBlockDomain/name[BlockDomain_id]+' are finished calculating');
				}
				container PerVertrekmoment := for_each_ne(Advanced/MeasureMoments/Name, 'PerDepartureMoment_T(' + string(id(Advanced/MeasureMoments)) + ')');
				Template PerDepartureMoment_T
				{
					parameter<Advanced/MeasureMoments> m;
					///
					// in the model: the Pareto-optimal options of this block and departure moment; here test data
					unit<uint64> Result_AfterSecondCondensation := range(uint64, uint64(0), uint64(1000 + 100 * m))
					{
						attribute<uint32> i              := uint32(id(.));
						attribute<Places> From_Place_rel := value(uint32(BlockDomain_id) * 25 + (i / 7) % 25, Places);
						attribute<Places> To_Place_rel   := value((i * i + 3 * uint32(m)) % 50, Places);
						attribute<bool>   IsDirectTravel := i % 3 == 0;
						attribute<Modes>  V_mode_rel     := value(uint8((i / 3) % 3), Modes);
						attribute<bool>   NeedsCar       := i % 5 == 0;
						attribute<bool>   NeedsBike      := i % 7 == 0;
					}
					container CreateExports
					{
						container LongFormat
						{
							container Results_Traveltime
							{
								// in the model: writes the CSV file of this block and departure moment
								parameter<string> Generate := 'ok';

								// container Tellingen: per departure moment, see below
							}
							container Results_Traveltime_Fence := PhaseContainer(Results_Traveltime, 'Results in '+NetworkSetup/Impl/OrgBlockDomain/name[BlockDomain_id]+' are finished calculating');
						}
					}
				}
			}
		}
	}
}
  • MeasureMoments are the departure moments. In the model they are chosen from a sample day. Their Name names the container of each moment; their TemplatableTextShrt names the columns of the files.
  • ConfigurationPerBlock makes a container per block of origins with the template ConfigurationPerOrgBlock_T, and in each block PerVertrekmoment makes a container per departure moment with PerDepartureMoment_T (Module 3).
  • Result_AfterSecondCondensation is the final result of one block and one departure moment: one element per option, with the relations From_Place_rel and To_Place_rel to Places, IsDirectTravel, the mode of a direct trip V_mode_rel, NeedsCar and NeedsBike. Here they are computed from the element number i, so that the counts are easy to check. Each block has its own 25 origins, as in the model.
  • Modes/V holds one parameter per mode (Modes/V/Walking and so on), made with for_each_nedv, to compare V_mode_rel with.
  • OD_code has one element per pair of places, (origin, destination); the counts use it for their pair keys. In the model it sits next to Places, so that all departure moments of a block share it.
  • In the model, Results_Traveltime/Generate writes the CSV file of a block and departure moment; here it only returns 'ok'.
  • The two PhaseContainers are the fences discussed below.

per departure moment

This Tellingen container belongs in Results_Traveltime. R is the final result of the block and departure moment:

container Tellingen
{
	unit<uint64>      R                   := Result_AfterSecondCondensation;
	attribute<bool>   IsKeten         (R) := NOT(R/IsDirectTravel);
	attribute<bool>   IsWW            (R) := R/IsDirectTravel && R/V_mode_rel == Modes/v/Walking;
	attribute<bool>   IsCC            (R) := R/IsDirectTravel && R/V_mode_rel == Modes/v/Cycling;
	attribute<bool>   IsAA            (R) := R/IsDirectTravel && R/V_mode_rel == Modes/v/Car;
	attribute<bool>   IsZonderAuto    (R) := NOT(R/NeedsCar);
	attribute<bool>   IsZonderAutoEnFiets (R) := NOT(R/NeedsCar) && NOT(R/NeedsBike);

	attribute<uint64> Rijen (Classifications/OD_Tellingen_Soort) := union_data(Classifications/OD_Tellingen_Soort
		, #R
		, sum_uint64(IsKeten)
		, sum_uint64(IsWW)
		, sum_uint64(IsCC)
		, sum_uint64(IsAA));

	attribute<OD_code> OD_key         (R) := combine_data(OD_code, R/From_Place_rel, R/To_Place_rel);
	unit<uint64> Paren := unique(OD_key)
	{
		attribute<bool> OV_keten             := any(../IsKeten, ../Paren_rel);
		attribute<bool> WW                   := any(../IsWW, ../Paren_rel);
		attribute<bool> CC                   := any(../IsCC, ../Paren_rel);
		attribute<bool> AA                   := any(../IsAA, ../Paren_rel);
		attribute<bool> zonder_auto          := any(../IsZonderAuto, ../Paren_rel);
		attribute<bool> zonder_auto_en_fiets := any(../IsZonderAutoEnFiets, ../Paren_rel);
	}
	attribute<Paren>  Paren_rel       (R) := rlookup(OD_key, Paren/values);
	attribute<uint64> Paren_Telling (Classifications/OD_Tellingen_Groep) := union_data(Classifications/OD_Tellingen_Groep
		, #Paren
		, sum_uint64(Paren/OV_keten)
		, sum_uint64(Paren/WW)
		, sum_uint64(Paren/CC)
		, sum_uint64(Paren/AA)
		, sum_uint64(Paren/zonder_auto)
		, sum_uint64(Paren/zonder_auto_en_fiets));
}

Things to notice:

  • The conditions are bool attributes on R, computed for all rows at once.
  • union_data(classification, ...) fills the table of the kinds (Rijen) and the table of the groups (Paren_Telling) with one value per element, in the order of the classification: the counterpart of a Python list. Each value is a parameter, and a parameter fills one element.
  • The first value is the total, #R or #Paren: the number of elements of a uint64 domain unit, so already a uint64. The others count the true values of a condition with sum_uint64, which sums a bool attribute straight into a uint64.
  • An OD pair becomes one element of OD_code, the combined domain of origins and destinations made with combine_unit_uint64(Places, Places). combine_data relates every row to the element of its pair, so OD_key is a relation, not a number computed by hand from two indices, and no conversion is needed. unique makes a new domain unit with one element per distinct pair; its attribute values holds the pair as an OD_code. rlookup relates each row to its pair, and any(condition, relation) tells per pair whether at least one of its rows meets the condition: the counterpart of adding a pair to a set under a condition.

per block: the union over the moments

This Tellingen container belongs in Traveltime_ForEachDepTime. The block takes the pairs of its departure moments together. union_unit_uint64 stacks the pair domains of the moments into one, union_data stacks their attributes, and unique and any do the rest, as before. The lists of the moments are generated with indirect expressions (='...') and asList over MeasureMoments, so the number of departure moments stays a model parameter (Module 3). Paren_Samen fills the group table with union_data again:

container Tellingen
{
	container Rijen := for_each_nedv(Advanced/MeasureMoments/Name
		, 'PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Rijen'
		, Classifications/OD_Tellingen_Soort, uint64);
	container Paren := for_each_nedv(Advanced/MeasureMoments/Name
		, 'PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren_Telling'
		, Classifications/OD_Tellingen_Groep, uint64);

	unit<uint64> AlleParen := ='union_unit_uint64('+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren', ',')+')'
	{
		attribute<OD_code> OD_key              := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/values', ',')+')';
		attribute<bool>   OV_keten             := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/OV_keten', ',')+')';
		attribute<bool>   WW                   := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/WW', ',')+')';
		attribute<bool>   CC                   := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/CC', ',')+')';
		attribute<bool>   AA                   := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/AA', ',')+')';
		attribute<bool>   zonder_auto          := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/zonder_auto', ',')+')';
		attribute<bool>   zonder_auto_en_fiets := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/zonder_auto_en_fiets', ',')+')';
		attribute<UniekeParen> Uniek_rel       := rlookup(OD_key, UniekeParen/values);
	}
	unit<uint64> UniekeParen := unique(AlleParen/OD_key)
	{
		attribute<bool> OV_keten             := any(AlleParen/OV_keten, AlleParen/Uniek_rel);
		attribute<bool> WW                   := any(AlleParen/WW, AlleParen/Uniek_rel);
		attribute<bool> CC                   := any(AlleParen/CC, AlleParen/Uniek_rel);
		attribute<bool> AA                   := any(AlleParen/AA, AlleParen/Uniek_rel);
		attribute<bool> zonder_auto          := any(AlleParen/zonder_auto, AlleParen/Uniek_rel);
		attribute<bool> zonder_auto_en_fiets := any(AlleParen/zonder_auto_en_fiets, AlleParen/Uniek_rel);
	}
	attribute<uint64> Paren_Samen (Classifications/OD_Tellingen_Groep) := union_data(Classifications/OD_Tellingen_Groep
		, #UniekeParen
		, sum_uint64(UniekeParen/OV_keten)
		, sum_uint64(UniekeParen/WW)
		, sum_uint64(UniekeParen/CC)
		, sum_uint64(UniekeParen/AA)
		, sum_uint64(UniekeParen/zonder_auto)
		, sum_uint64(UniekeParen/zonder_auto_en_fiets));
}

over all blocks: add up and write

This container belongs in ConfigurationPerBlock/Generate_Output. The blocks are added up with an indirect expression that writes out the sum over all blocks (Block_1of450/... + Block_2of450/... + ...). The two files are plain parameter<string> items with StorageType = "str", which writes the string as a text file:

container OD_Tellingen
{
	unit<uint8>  Soort   := Classifications/OD_Tellingen_Soort;
	unit<uint8>  Groep   := Classifications/OD_Tellingen_Groep;
	unit<uint32> Moments := /ModelParameters/Advanced/MeasureMoments;

	container PerMoment := for_each_ne(Moments/Name, 'Moment_T(' + quote(Moments/Name) + ')');
	Template Moment_T
	{
		parameter<string> Moment;
		///
		attribute<uint64> Rijen (Soort) := =asList(/NetworkSetup/impl/OrgBlockDomain/name + '/PublicTransport/Generate_Output/Traveltime_ForEachDepTime_Fence/Tellingen/Rijen/' + Moment, ' + ');
		attribute<uint64> Paren (Groep) := =asList(/NetworkSetup/impl/OrgBlockDomain/name + '/PublicTransport/Generate_Output/Traveltime_ForEachDepTime_Fence/Tellingen/Paren/' + Moment, ' + ');
	}
	attribute<uint64> Rijen_Totaal (Soort) := =asList('PerMoment/' + Moments/Name + '/Rijen', ' + ');
	attribute<uint64> Paren_Samen  (Groep) := =asList(/NetworkSetup/impl/OrgBlockDomain/name + '/PublicTransport/Generate_Output/Traveltime_ForEachDepTime_Fence/Tellingen/Paren_Samen', ' + ');

	parameter<string> MomentKop := asList(Moments/TemplatableTextShrt, ';');
	attribute<string> RijenRegel (Soort) := ='Soort/label + ' + asList(''';'' + string(PerMoment/' + Moments/Name + '/Rijen)', ' + ') + ' + '';'' + string(Rijen_Totaal)';
	attribute<string> ParenRegel (Groep) := ='Groep/label + ' + asList(''';'' + string(PerMoment/' + Moments/Name + '/Paren)', ' + ') + ' + '';'' + string(Paren_Samen)';
	parameter<string> File_Rijen := 'soort;' + MomentKop + ';totaal\n' + asList(RijenRegel, '\n') + '\n'
	,	StorageName = "=/ModelParameters/Advanced/OutputDir + '/od_tellingen.csv'", StorageType = "str";
	parameter<string> File_Paren := 'groep;' + MomentKop + ';samen\n' + asList(ParenRegel, '\n') + '\n'
	,	StorageName = "=/ModelParameters/Advanced/OutputDir + '/unieke_od_tellingen.csv'", StorageType = "str";
	parameter<string> Generate := File_Rijen + File_Paren;
}

one request, behind a fence

The output step computes 450 blocks with four departure moments each; the full results of all of them would never fit in memory together. The model therefore puts the results of each departure moment and of each block behind a fence, a PhaseContainer: once a phase is finished, its large intermediate results can be freed, and only what the request still needs is kept. Here that is small: per moment the five row counts and the pairs (about 32 thousand per block) with six flags, per block twelve numbers per moment.

The intermediate results behind the fences exist only while the blocks are being computed. So the counts are requested together with the CSV files, in one item (it is in the skeleton above):

parameter<string> OUTPUT_Generate_PublicTransport_fullOD_long_CSVFiles :=
	=asList(NetworkSetup/Impl/OrgBlockDomain/name+'/PublicTransport/Generate_Output/Traveltime_ForEachDepTime_Fence/Generate',' + ')
	+ ' + OD_Tellingen/Generate';

This is also the price of the GeoDMS implementation: it cannot count output that already exists. For that it needs the results in memory, so the whole output step (2.5 hours) would have to run again. That is why the Python script exists next to it.

translating the idioms

Python (OdTellingen.py) GeoDMS
a row of a CSV file an element of the domain unit R
soort = 'OV-keten' if '_' in mode else mode the conditions IsKeten, IsWW, IsCC and IsAA
a Counter of the kinds sum_uint64(IsKeten) and so on, one per kind
od = (row[0], row[1]), a tuple of two strings OD_key := combine_data(OD_code, From_Place_rel, To_Place_rel), a relation to the combined domain OD_code
a set of pairs unique(OD_key), a new domain unit
per['zonder auto'].add(od) under a condition any(IsZonderAuto, Paren_rel)
samen[g] \|= s over the moments union_unit_uint64 and union_data, then unique and any
Pool(8).imap_unordered(blok, ...) over the blocks the GeoDMS’s own multithreading, blocks generated with for_each
R.update(r) over the blocks an indirect expression asList(..., ' + ')
the lists SOORTEN and GROEPEN, and a count per element in that order the classifications and union_data(classification, ...)
f.write(...) a parameter<string> with StorageType = "str"

performance

how it was measured

All measurements were made on 4 October 2026 on one workstation: AMD Ryzen 9 7950X (16 cores, 32 threads), 128 GB RAM, Windows 11, Python 3.14.5 and GeoDMS 20.22.1. The machine was busy with another run of the same model at the same time, so the absolute times vary with the load (the average CPU load during the measurements was 28 to 96%, including the measured process). The CSV files were in the operating system’s file cache, so reading them cost almost no disk time.

  • Python: time.perf_counter (wall time) and time.process_time (CPU time) around the function blok of the script, for each block in the worker processes and for the whole run in the main process, and psutil for the peak memory of the workers. To separate reading and parsing from counting, the same files were also only read (bytes), and only parsed (csv.reader, without counting).
  • GeoDMS: the log file of a GeoDmsRun call (option /L) contains a line per operation with the item and its duration, for example

    2026-10-04 13:33:48[17][.][performance]oper unique [[/NetworkSetup/ConfigurationPerBlock/Block_1of450/.../Tellingen/Paren]]: 3.8ms
    

    The cost of the counts is the sum over the operations on items under .../Tellingen/, plus the unnamed intermediate results inside their expressions, such as the pcount in Rijen and the sums in Paren_Telling. Block 1 of the morning of 6 October was computed this way three times: once for the counts alone, twice together with its four CSV files. These runs used the first version of the code, which counted the rows per kind with pcount over a relation from the rows to the kinds, picked the count per group with switch and case on the element number, and computed the pair key by hand as origin times #Places plus destination; the code on this page gathers the same counts with union_data and sum_uint64, and makes the pair key with combine_data. Block 1 has 393,886 rows over the four moments, close to the average of 384,445 rows per block. The Python measurements of block 1 were repeated eight times.

one block

Block 1, four departure moments, 393,886 rows, in seconds; the fastest and the median run:

step Python, one process (8 runs) GeoDMS, inside the output step (3 runs)
read the four files (53 MB, from the file cache) 0.02 / 0.03 –
parse the CSV text 0.41 / 0.57 –
the counts themselves 0.32 / 0.46 0.053 / 0.054
total 0.73 / 1.03 0.053 / 0.054

The Python times of the counts are the full script minus parsing alone. The slowest Python run took 1.71 s, at a machine load of 96%.

The GeoDMS times per departure moment were 7 to 14 ms, of which unique over the 98 thousand keys took 3.5 to 7 ms, and rlookup, pcount and each any 0.3 to 3 ms. The union over the moments (128 thousand pair rows into 32 thousand pairs) took 8 to 12 ms, except in one run, in which a single unique took 254 ms and the block 0.36 s in total; with another run competing for the machine, an operation can wait for memory or a thread.

For scale: the whole output of block 1, from reading the chain stores to writing the four CSV files, took 8.5 to 11 minutes, or 573 to 893 seconds summed over the operations. The counts are less than 0.01% of that (0.06% in the slow run).

the whole morning

All 450 blocks, 1,800 files, 23.3 GB, 173,000,050 rows:

  Python, afterwards GeoDMS, in the output step
input 1,800 CSV files, read and parsed none: the results are in memory
wall time 65 s, with 8 processes not separately visible: the output step without the counts took 2 h 36 min
CPU time 491 s about 25 s (450 blocks x 54 ms), so at most 25 s of the 2 h 36 min
extra memory at most 61 MB per process about 2 MB per departure moment of a block

where the difference comes from

  1. Reading and parsing. Of each 31-column row, Python makes new string objects for every field and uses five. That was about 35 to 55% of its time. The GeoDMS reads nothing: the counts use arrays the output step already holds.
  2. Types. Python identifies an OD pair by a tuple of two strings, which it allocates, hashes and stores per row in up to four sets. In the GeoDMS a pair is one element of the combined domain OD_code, stored as one 8-byte integer, and unique over 98 thousand integers takes a few milliseconds.
  3. Interpretation. The Python loop executes a dozen interpreted operations per row; a GeoDMS operator is one compiled loop over a typed array. For the counting itself, the GeoDMS was 6 to 9 times faster on the same rows; including the parsing that it does not need, 14 to 19 times.
  4. Parallelism at a different grain. Python runs independent blocks in separate processes; the GeoDMS runs operations, and segments of large operations, in threads. The output step keeps all cores busy anyway, so the counts add no wall time worth mentioning.

what to take from this

  • Count where the data is. When the data is in memory anyway, a summary in the model costs a few dozen lines of configuration and, here, well under 1% of the output step. It is made in every run, it is always consistent with the output, and it is traceable like any other item.
  • A short script is right for data that already exists. One minute on 8 cores for 173 million rows of text is fine for a one-off or retrospective check. Stream the rows, keep the memory per worker small, and run independent files in parallel. When you have to read the same output often, a binary column format such as parquet (Module 2D, GeoDMS and Python) or a native GeoDMS format (Module 2C, GeoDMS own data formats) avoids most of the parsing.
  • Two implementations are also a test. The Python script and the GeoDMS code were written separately from the same definitions. For block 1 they gave identical counts, which checks both.

try it yourself!

  1. Assemble the example of this page in one file: the skeleton, with the two classifications in Classifications and the three Tellingen containers at the places marked see above and see below. Open it in the GeoDMS GUI and update /NetworkSetup/ConfigurationPerBlock/Generate_Output/OUTPUT_Generate_PublicTransport_fullOD_long_CSVFiles, or run that item with GeoDmsRun. The folder %LocalDataProjDir%/Output then has od_tellingen.csv and unieke_od_tellingen.csv, together:

     soort;07h00m;07h15m;totaal
     alle;2000;2200;4200
     OV-keten;1332;1466;2798
     WW;224;246;470
     CC;222;244;466
     AA;222;244;466
     groep;07h00m;07h15m;samen
     alle;584;584;1060
     OV-keten;584;584;1060
     WW;212;230;416
     CC;208;226;420
     AA;204;224;408
     zonder auto;484;484;860
     zonder auto en fiets;424;424;768
    

    Then change the test data, for example NeedsCar := i % 2 == 0, predict which numbers change, and run it again.

  2. Make a large table yourself. In Python, write a CSV file with 10 million rows and two columns, org and dest, filled with names such as org123 and dest4567 (draw the numbers at random below 10,000). Count the unique (org, dest) pairs with a set of tuples of strings, and time it with time.perf_counter. Then turn the names into integers first (a dict per column) and count the unique keys org * 10000 + dest. How much of the time went into the tuples of strings, and how much into reading the file?
  3. In the GeoDMS, read the same file (Module 2A). Make domain units of the unique origins and destinations with unique, relate the rows to them with rlookup, combine the two relations into one with combine_unit_uint64 and combine_data, and count the elements of unique on it. Compare the count with your Python result, and look up the time of each step in View > Calculation times (Deepening II).
  4. Add a condition, for example the destination number is even, as a bool attribute, and count with any(condition, relation) the pairs that meet it at least once. Write both counts to a text file with a parameter<string> and StorageType = "str".
  5. Think about a summary in your own model: is the data it needs in memory during a run anyway? Then which would you choose, a few lines in the model or a script afterwards?

Go to previous deepening: Deepening II, Performance and memory