Deepening III, Python and GeoDMS side by side

learning objective: reading the same aggregation in Python and in the GeoDMS, mapping the Python idioms onto GeoDMS operators, and knowing what makes each implementation fast or slow
introduction
Module 2D, GeoDMS and Python shows how the GeoDMS and Python exchange data. This deepening looks at the two from another angle: one calculation, written once in Python and once in the GeoDMS, run on real model output of 173 million rows. You will see how the idioms of one language translate into the other, and what the difference in speed is made of.
The case comes from NetworkModel_PBL, an accessibility model made for the Netherlands Environmental Assessment Agency (PBL). For every origin-destination pair (OD pair) it calculates the Pareto-optimal travel options: options for which no alternative is better on all criteria, such as travel time and price. An option is a public transport chain with pre- and post-transport, or a direct trip on foot, by bike or by car. The output step writes one CSV file per block of origins (450 blocks) and per departure moment (07:00, 07:15, 07:30 and 07:45), with one row per option. For the morning of 6 October 2026 that is 1,800 files, 23.3 GB, and 173,000,050 rows.
Both implementations are in use. The GeoDMS code is part of the output step and makes the counts in every run; the Python script makes the same counts for output that was written before the GeoDMS code existed.
This page assumes you know relations and aggregations (Module 1c, Learning the basic concepts of GeoDMS, calculations over multiple domains), for_each and indirect expressions (Module 3, Meta scripting ‐ templates and for_each), and the performance concepts of Deepening II, Performance and memory.
the question
Each row of an output file has, among its 31 columns, the origin (OrgName), the destination (DestName), the modes used (ModeUsed_<moment>), and whether the option needs a car or a bike (NeedsCar_<moment>, NeedsBike_<moment>, 0 or 1):
OrgName;DestName;Traveltime_07h00m;Price_07h00m;Price_Augm_07h00m;ModeUsed_07h00m;NeedsCar_07h00m;NeedsBike_07h00m;...
PandBlok_65575;Rheden;40.133335;5.21;8.23;W_BBR_W;0;0;...
ModeUsed is WW, CC or AA for a direct trip on foot, by bike or by car. A public transport chain has underscores, as in W_BBR_W: walk to the first stop, three public transport legs, walk from the last stop.
Two small tables summarise such output at a glance. The model’s users are Dutch, and so are the labels:
od_tellingen.csv(counts of OD rows): the number of rows per kind of option, per departure moment and in total. The kinds arealle(all),OV-keten(a public transport chain),WW,CCandAA.unieke_od_tellingen.csv(counts of unique OD pairs): the number of unique OD pairs with at least one option of that kind, or with at least one option for a group of travellers:zonder auto(without a car:NeedsCaris 0) andzonder auto en fiets(without a car and a bike). Per departure moment, and for the four moments together (samen).
For the morning of 6 October 2026:
| soort (kind) | 07h00m | 07h15m | 07h30m | 07h45m | totaal |
|---|---|---|---|---|---|
| alle | 43,214,954 | 43,220,795 | 43,320,686 | 43,243,615 | 173,000,050 |
| OV-keten | 3,821,320 | 3,826,076 | 3,941,686 | 3,854,368 | 15,443,450 |
| WW | 107,284 | 107,284 | 107,284 | 107,284 | 429,136 |
| CC | 836,827 | 836,827 | 836,829 | 836,828 | 3,347,311 |
| AA | 38,449,523 | 38,450,608 | 38,434,887 | 38,445,135 | 153,780,153 |
| groep (group) | 07h00m | 07h15m | 07h30m | 07h45m | samen |
|---|---|---|---|---|---|
| alle | 14,443,660 | 14,444,579 | 14,443,642 | 14,445,007 | 14,448,616 |
| OV-keten | 1,533,524 | 1,539,869 | 1,580,944 | 1,549,725 | 2,087,775 |
| WW | 107,284 | 107,284 | 107,284 | 107,284 | 107,284 |
| CC | 836,827 | 836,827 | 836,829 | 836,828 | 836,830 |
| AA | 14,333,586 | 14,333,004 | 14,331,709 | 14,331,660 | 14,387,855 |
| zonder auto | 2,040,917 | 2,040,659 | 2,078,823 | 2,048,502 | 2,516,369 |
| zonder auto en fiets | 1,589,215 | 1,594,529 | 1,635,229 | 1,603,859 | 2,132,806 |
So 173 million options connect 14.4 million OD pairs, and 2.5 million of those pairs can be travelled without a car at one of the four moments.
The code below keeps the Dutch names of the model. A small glossary: tellingen = counts, rijen = rows, paren = pairs, soort = kind, groep = group, keten = chain, OV = public transport, blok = block, moment = departure moment, uitvoermap = output folder.
One property of the data makes both implementations simple: every origin belongs to exactly one block, so every OD pair occurs in one block only. The counts of the blocks can just be added up. Only the union over the four departure moments needs care, and that union stays within a block.
the Python implementation
The script handles one block at a time. It streams the rows of the four files of the block through csv.reader, counts the rows per kind in a Counter, and collects the pairs per group in Python sets of (origin, destination) tuples. The union over the moments is a set union. Because the blocks are independent, a multiprocessing.Pool spreads them over 8 processes, and the main process adds up the results:
import csv, glob, os, re
from collections import Counter, defaultdict
from multiprocessing import Pool
NAME_RE = re.compile(r'tt_\d{8}_(\d\dh\d\dm)_(\d+)of(\d+)\.csv$')
def blok(files):
rijen, paren, samen = Counter(), Counter(), defaultdict(set)
for p in files:
m = NAME_RE.search(os.path.basename(p)).group(1)
per = defaultdict(set)
with open(p, newline='', encoding='utf-8') as fh:
r = csv.reader(fh, delimiter=';')
h = next(r)
im, ic, ib = h.index('ModeUsed_' + m), h.index('NeedsCar_' + m), h.index('NeedsBike_' + m)
for row in r:
if not row:
continue
od = (row[0], row[1])
mode = row[im]
soort = 'OV-keten' if '_' in mode else mode
rijen[(m, 'alle')] += 1
rijen[(m, soort)] += 1
per['alle'].add(od)
per[soort].add(od)
if row[ic] == '0':
per['zonder auto'].add(od)
if row[ib] == '0':
per['zonder auto en fiets'].add(od)
for g, s in per.items():
paren[(m, g)] += len(s)
samen[g] |= s
return rijen, paren, Counter({g: len(s) for g, s in samen.items()})
# in main(): group the files per block, count the blocks in parallel, add up
per_blok = defaultdict(list)
for p in glob.glob(os.path.join(uitvoermap, 'PerBlock', 'tt_*.csv')):
per_blok[NAME_RE.search(os.path.basename(p)).group(2)].append(p)
R, P, S = Counter(), Counter(), Counter()
with Pool(8) as pool:
for r, p, s in pool.imap_unordered(blok, per_blok.values()):
R.update(r); P.update(p); S.update(s)
# ... and write R, P and S as od_tellingen.csv and unieke_od_tellingen.csv
The full script, which also refuses to count when a block or a moment is missing, is batch/OdTellingen.py.
Note how much of this code has nothing to do with the question itself: finding the files by name, finding the columns by header, parsing 31 text columns per row to use 5 of them, and identifying an OD pair by two strings.
the GeoDMS implementation
In the GeoDMS the counts are made in the output step itself, from the final result of each departure moment. That result is in memory anyway, because the CSV file is written from it. The code has three levels: per departure moment, per block, and over all blocks.
Together, the code blocks of this section form one complete configuration. Put the two classifications and the three Tellingen containers in the skeleton below, at the places marked see above and see below, and it runs as it is, on small test data instead of the model’s results (see try it yourself!).
the classifications and the model around the counts
The kinds and the groups are two small classifications. The code uses their order and element names; the files show their labels:
unit<uint8> OD_Tellingen_Soort : nrofrows = 5
{
attribute<string> name : ['alle', 'OV_keten', 'WW', 'CC', 'AA'];
attribute<string> label : ['alle', 'OV-keten', 'WW', 'CC', 'AA'];
}
unit<uint8> OD_Tellingen_Groep : nrofrows = 7
{
attribute<string> name : ['alle', 'OV_keten', 'WW', 'CC', 'AA', 'zonder_auto', 'zonder_auto_en_fiets'];
attribute<string> label : ['alle', 'OV-keten', 'WW', 'CC', 'AA', 'zonder auto', 'zonder auto en fiets'];
}
The counts sit in the container structure of the model. The skeleton below has that structure, and declares the items the counts use, with test data where the model has its own calculations:
container Tellingen_Example
{
container Classifications
{
// unit<uint8> OD_Tellingen_Soort and unit<uint8> OD_Tellingen_Groep: see above
}
container ModelParameters
{
container Advanced
{
// in the model: the departure moments chosen from the sample day
unit<uint32> MeasureMoments : nrofrows = 2
{
attribute<string> Name : ['At_07h00m00s', 'At_07h15m00s'];
attribute<string> TemplatableTextShrt : ['07h00m', '07h15m'];
}
parameter<string> OutputDir := '%LocalDataProjDir%/Output';
}
}
container NetworkSetup : using = "Classifications;ModelParameters"
{
container impl
{
// in the model: 450 blocks of origins
unit<uint32> OrgBlockDomain : nrofrows = 2
{
attribute<string> name : ['Block_1of2', 'Block_2of2'];
}
}
container ConfigurationPerBlock := for_each_ne(impl/OrgBlockDomain/name, 'ConfigurationPerOrgBlock_T(' + string(id(impl/OrgBlockDomain)) + ')')
{
container Generate_Output
{
parameter<string> OUTPUT_Generate_PublicTransport_fullOD_long_CSVFiles := =asList(NetworkSetup/Impl/OrgBlockDomain/name+'/PublicTransport/Generate_Output/Traveltime_ForEachDepTime_Fence/Generate',' + ') + ' + OD_Tellingen/Generate';
// container OD_Tellingen: over all blocks, see below
}
}
Template ConfigurationPerOrgBlock_T
{
parameter<NetworkSetup/impl/OrgBlockDomain> BlockDomain_id;
///
container PublicTransport
{
// in the model: all stops, origins and destinations, and the modes of the feed plus walking, cycling and car
unit<uint32> Places : nrofrows = 50;
// one element per (origin, destination) pair
unit<uint64> OD_code := combine_unit_uint64(Places, Places);
unit<uint8> Modes : nrofrows = 3
{
attribute<string> name : ['Walking', 'Cycling', 'Car'];
container V := for_each_nedv(name, String(ID(.))+'[..]', void, .);
}
container Generate_Output
{
container Traveltime_ForEachDepTime
{
parameter<string> Generate := ='add('+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Generate',',')+')';
// container Tellingen: per block, see below
}
container Traveltime_ForEachDepTime_Fence := PhaseContainer(Traveltime_ForEachDepTime, 'Results for all departure times in '+NetworkSetup/Impl/OrgBlockDomain/name[BlockDomain_id]+' are finished calculating');
}
container PerVertrekmoment := for_each_ne(Advanced/MeasureMoments/Name, 'PerDepartureMoment_T(' + string(id(Advanced/MeasureMoments)) + ')');
Template PerDepartureMoment_T
{
parameter<Advanced/MeasureMoments> m;
///
// in the model: the Pareto-optimal options of this block and departure moment; here test data
unit<uint64> Result_AfterSecondCondensation := range(uint64, uint64(0), uint64(1000 + 100 * m))
{
attribute<uint32> i := uint32(id(.));
attribute<Places> From_Place_rel := value(uint32(BlockDomain_id) * 25 + (i / 7) % 25, Places);
attribute<Places> To_Place_rel := value((i * i + 3 * uint32(m)) % 50, Places);
attribute<bool> IsDirectTravel := i % 3 == 0;
attribute<Modes> V_mode_rel := value(uint8((i / 3) % 3), Modes);
attribute<bool> NeedsCar := i % 5 == 0;
attribute<bool> NeedsBike := i % 7 == 0;
}
container CreateExports
{
container LongFormat
{
container Results_Traveltime
{
// in the model: writes the CSV file of this block and departure moment
parameter<string> Generate := 'ok';
// container Tellingen: per departure moment, see below
}
container Results_Traveltime_Fence := PhaseContainer(Results_Traveltime, 'Results in '+NetworkSetup/Impl/OrgBlockDomain/name[BlockDomain_id]+' are finished calculating');
}
}
}
}
}
}
}
MeasureMomentsare the departure moments. In the model they are chosen from a sample day. TheirNamenames the container of each moment; theirTemplatableTextShrtnames the columns of the files.ConfigurationPerBlockmakes a container per block of origins with the templateConfigurationPerOrgBlock_T, and in each blockPerVertrekmomentmakes a container per departure moment withPerDepartureMoment_T(Module 3).Result_AfterSecondCondensationis the final result of one block and one departure moment: one element per option, with the relationsFrom_Place_relandTo_Place_reltoPlaces,IsDirectTravel, the mode of a direct tripV_mode_rel,NeedsCarandNeedsBike. Here they are computed from the element numberi, so that the counts are easy to check. Each block has its own 25 origins, as in the model.Modes/Vholds one parameter per mode (Modes/V/Walkingand so on), made withfor_each_nedv, to compareV_mode_relwith.OD_codehas one element per pair of places, (origin, destination); the counts use it for their pair keys. In the model it sits next toPlaces, so that all departure moments of a block share it.- In the model,
Results_Traveltime/Generatewrites the CSV file of a block and departure moment; here it only returns'ok'. - The two
PhaseContainers are the fences discussed below.
per departure moment
This Tellingen container belongs in Results_Traveltime. R is the final result of the block and departure moment:
container Tellingen
{
unit<uint64> R := Result_AfterSecondCondensation;
attribute<bool> IsKeten (R) := NOT(R/IsDirectTravel);
attribute<bool> IsWW (R) := R/IsDirectTravel && R/V_mode_rel == Modes/v/Walking;
attribute<bool> IsCC (R) := R/IsDirectTravel && R/V_mode_rel == Modes/v/Cycling;
attribute<bool> IsAA (R) := R/IsDirectTravel && R/V_mode_rel == Modes/v/Car;
attribute<bool> IsZonderAuto (R) := NOT(R/NeedsCar);
attribute<bool> IsZonderAutoEnFiets (R) := NOT(R/NeedsCar) && NOT(R/NeedsBike);
attribute<uint64> Rijen (Classifications/OD_Tellingen_Soort) := union_data(Classifications/OD_Tellingen_Soort
, #R
, sum_uint64(IsKeten)
, sum_uint64(IsWW)
, sum_uint64(IsCC)
, sum_uint64(IsAA));
attribute<OD_code> OD_key (R) := combine_data(OD_code, R/From_Place_rel, R/To_Place_rel);
unit<uint64> Paren := unique(OD_key)
{
attribute<bool> OV_keten := any(../IsKeten, ../Paren_rel);
attribute<bool> WW := any(../IsWW, ../Paren_rel);
attribute<bool> CC := any(../IsCC, ../Paren_rel);
attribute<bool> AA := any(../IsAA, ../Paren_rel);
attribute<bool> zonder_auto := any(../IsZonderAuto, ../Paren_rel);
attribute<bool> zonder_auto_en_fiets := any(../IsZonderAutoEnFiets, ../Paren_rel);
}
attribute<Paren> Paren_rel (R) := rlookup(OD_key, Paren/values);
attribute<uint64> Paren_Telling (Classifications/OD_Tellingen_Groep) := union_data(Classifications/OD_Tellingen_Groep
, #Paren
, sum_uint64(Paren/OV_keten)
, sum_uint64(Paren/WW)
, sum_uint64(Paren/CC)
, sum_uint64(Paren/AA)
, sum_uint64(Paren/zonder_auto)
, sum_uint64(Paren/zonder_auto_en_fiets));
}
Things to notice:
- The conditions are
boolattributes onR, computed for all rows at once. - union_data
(classification, ...)fills the table of the kinds (Rijen) and the table of the groups (Paren_Telling) with one value per element, in the order of the classification: the counterpart of a Python list. Each value is a parameter, and a parameter fills one element. - The first value is the total,
#Ror#Paren: the number of elements of auint64domain unit, so already auint64. The others count thetruevalues of a condition withsum_uint64, which sums aboolattribute straight into auint64. - An OD pair becomes one element of
OD_code, the combined domain of origins and destinations made with combine_unit_uint64(Places, Places). combine_data relates every row to the element of its pair, soOD_keyis a relation, not a number computed by hand from two indices, and no conversion is needed. unique makes a new domain unit with one element per distinct pair; its attributevaluesholds the pair as anOD_code. rlookup relates each row to its pair, and any(condition, relation)tells per pair whether at least one of its rows meets the condition: the counterpart of adding a pair to a set under a condition.
per block: the union over the moments
This Tellingen container belongs in Traveltime_ForEachDepTime. The block takes the pairs of its departure moments together. union_unit_uint64 stacks the pair domains of the moments into one, union_data stacks their attributes, and unique and any do the rest, as before. The lists of the moments are generated with indirect expressions (='...') and asList over MeasureMoments, so the number of departure moments stays a model parameter (Module 3). Paren_Samen fills the group table with union_data again:
container Tellingen
{
container Rijen := for_each_nedv(Advanced/MeasureMoments/Name
, 'PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Rijen'
, Classifications/OD_Tellingen_Soort, uint64);
container Paren := for_each_nedv(Advanced/MeasureMoments/Name
, 'PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren_Telling'
, Classifications/OD_Tellingen_Groep, uint64);
unit<uint64> AlleParen := ='union_unit_uint64('+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren', ',')+')'
{
attribute<OD_code> OD_key := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/values', ',')+')';
attribute<bool> OV_keten := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/OV_keten', ',')+')';
attribute<bool> WW := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/WW', ',')+')';
attribute<bool> CC := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/CC', ',')+')';
attribute<bool> AA := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/AA', ',')+')';
attribute<bool> zonder_auto := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/zonder_auto', ',')+')';
attribute<bool> zonder_auto_en_fiets := ='union_data(., '+asList('PerVertrekmoment/'+Advanced/MeasureMoments/Name+'/CreateExports/LongFormat/Results_Traveltime_Fence/Tellingen/Paren/zonder_auto_en_fiets', ',')+')';
attribute<UniekeParen> Uniek_rel := rlookup(OD_key, UniekeParen/values);
}
unit<uint64> UniekeParen := unique(AlleParen/OD_key)
{
attribute<bool> OV_keten := any(AlleParen/OV_keten, AlleParen/Uniek_rel);
attribute<bool> WW := any(AlleParen/WW, AlleParen/Uniek_rel);
attribute<bool> CC := any(AlleParen/CC, AlleParen/Uniek_rel);
attribute<bool> AA := any(AlleParen/AA, AlleParen/Uniek_rel);
attribute<bool> zonder_auto := any(AlleParen/zonder_auto, AlleParen/Uniek_rel);
attribute<bool> zonder_auto_en_fiets := any(AlleParen/zonder_auto_en_fiets, AlleParen/Uniek_rel);
}
attribute<uint64> Paren_Samen (Classifications/OD_Tellingen_Groep) := union_data(Classifications/OD_Tellingen_Groep
, #UniekeParen
, sum_uint64(UniekeParen/OV_keten)
, sum_uint64(UniekeParen/WW)
, sum_uint64(UniekeParen/CC)
, sum_uint64(UniekeParen/AA)
, sum_uint64(UniekeParen/zonder_auto)
, sum_uint64(UniekeParen/zonder_auto_en_fiets));
}
over all blocks: add up and write
This container belongs in ConfigurationPerBlock/Generate_Output. The blocks are added up with an indirect expression that writes out the sum over all blocks (Block_1of450/... + Block_2of450/... + ...). The two files are plain parameter<string> items with StorageType = "str", which writes the string as a text file:
container OD_Tellingen
{
unit<uint8> Soort := Classifications/OD_Tellingen_Soort;
unit<uint8> Groep := Classifications/OD_Tellingen_Groep;
unit<uint32> Moments := /ModelParameters/Advanced/MeasureMoments;
container PerMoment := for_each_ne(Moments/Name, 'Moment_T(' + quote(Moments/Name) + ')');
Template Moment_T
{
parameter<string> Moment;
///
attribute<uint64> Rijen (Soort) := =asList(/NetworkSetup/impl/OrgBlockDomain/name + '/PublicTransport/Generate_Output/Traveltime_ForEachDepTime_Fence/Tellingen/Rijen/' + Moment, ' + ');
attribute<uint64> Paren (Groep) := =asList(/NetworkSetup/impl/OrgBlockDomain/name + '/PublicTransport/Generate_Output/Traveltime_ForEachDepTime_Fence/Tellingen/Paren/' + Moment, ' + ');
}
attribute<uint64> Rijen_Totaal (Soort) := =asList('PerMoment/' + Moments/Name + '/Rijen', ' + ');
attribute<uint64> Paren_Samen (Groep) := =asList(/NetworkSetup/impl/OrgBlockDomain/name + '/PublicTransport/Generate_Output/Traveltime_ForEachDepTime_Fence/Tellingen/Paren_Samen', ' + ');
parameter<string> MomentKop := asList(Moments/TemplatableTextShrt, ';');
attribute<string> RijenRegel (Soort) := ='Soort/label + ' + asList(''';'' + string(PerMoment/' + Moments/Name + '/Rijen)', ' + ') + ' + '';'' + string(Rijen_Totaal)';
attribute<string> ParenRegel (Groep) := ='Groep/label + ' + asList(''';'' + string(PerMoment/' + Moments/Name + '/Paren)', ' + ') + ' + '';'' + string(Paren_Samen)';
parameter<string> File_Rijen := 'soort;' + MomentKop + ';totaal\n' + asList(RijenRegel, '\n') + '\n'
, StorageName = "=/ModelParameters/Advanced/OutputDir + '/od_tellingen.csv'", StorageType = "str";
parameter<string> File_Paren := 'groep;' + MomentKop + ';samen\n' + asList(ParenRegel, '\n') + '\n'
, StorageName = "=/ModelParameters/Advanced/OutputDir + '/unieke_od_tellingen.csv'", StorageType = "str";
parameter<string> Generate := File_Rijen + File_Paren;
}
one request, behind a fence
The output step computes 450 blocks with four departure moments each; the full results of all of them would never fit in memory together. The model therefore puts the results of each departure moment and of each block behind a fence, a PhaseContainer: once a phase is finished, its large intermediate results can be freed, and only what the request still needs is kept. Here that is small: per moment the five row counts and the pairs (about 32 thousand per block) with six flags, per block twelve numbers per moment.
The intermediate results behind the fences exist only while the blocks are being computed. So the counts are requested together with the CSV files, in one item (it is in the skeleton above):
parameter<string> OUTPUT_Generate_PublicTransport_fullOD_long_CSVFiles :=
=asList(NetworkSetup/Impl/OrgBlockDomain/name+'/PublicTransport/Generate_Output/Traveltime_ForEachDepTime_Fence/Generate',' + ')
+ ' + OD_Tellingen/Generate';
This is also the price of the GeoDMS implementation: it cannot count output that already exists. For that it needs the results in memory, so the whole output step (2.5 hours) would have to run again. That is why the Python script exists next to it.
translating the idioms
Python (OdTellingen.py) | GeoDMS |
|---|---|
| a row of a CSV file | an element of the domain unit R |
soort = 'OV-keten' if '_' in mode else mode | the conditions IsKeten, IsWW, IsCC and IsAA |
a Counter of the kinds | sum_uint64(IsKeten) and so on, one per kind |
od = (row[0], row[1]), a tuple of two strings | OD_key := combine_data(OD_code, From_Place_rel, To_Place_rel), a relation to the combined domain OD_code |
a set of pairs | unique(OD_key), a new domain unit |
per['zonder auto'].add(od) under a condition | any(IsZonderAuto, Paren_rel) |
samen[g] \|= s over the moments | union_unit_uint64 and union_data, then unique and any |
Pool(8).imap_unordered(blok, ...) over the blocks | the GeoDMS’s own multithreading, blocks generated with for_each |
R.update(r) over the blocks | an indirect expression asList(..., ' + ') |
the lists SOORTEN and GROEPEN, and a count per element in that order | the classifications and union_data(classification, ...) |
f.write(...) | a parameter<string> with StorageType = "str" |
performance
how it was measured
All measurements were made on 4 October 2026 on one workstation: AMD Ryzen 9 7950X (16 cores, 32 threads), 128 GB RAM, Windows 11, Python 3.14.5 and GeoDMS 20.22.1. The machine was busy with another run of the same model at the same time, so the absolute times vary with the load (the average CPU load during the measurements was 28 to 96%, including the measured process). The CSV files were in the operating system’s file cache, so reading them cost almost no disk time.
- Python:
time.perf_counter(wall time) andtime.process_time(CPU time) around the functionblokof the script, for each block in the worker processes and for the whole run in the main process, andpsutilfor the peak memory of the workers. To separate reading and parsing from counting, the same files were also only read (bytes), and only parsed (csv.reader, without counting). -
GeoDMS: the log file of a
GeoDmsRuncall (option/L) contains a line per operation with the item and its duration, for example2026-10-04 13:33:48[17][.][performance]oper unique [[/NetworkSetup/ConfigurationPerBlock/Block_1of450/.../Tellingen/Paren]]: 3.8msThe cost of the counts is the sum over the operations on items under
.../Tellingen/, plus the unnamed intermediate results inside their expressions, such as thepcountinRijenand thesums inParen_Telling. Block 1 of the morning of 6 October was computed this way three times: once for the counts alone, twice together with its four CSV files. These runs used the first version of the code, which counted the rows per kind withpcountover a relation from the rows to the kinds, picked the count per group withswitchandcaseon the element number, and computed the pair key by hand as origin times#Placesplus destination; the code on this page gathers the same counts withunion_dataandsum_uint64, and makes the pair key withcombine_data. Block 1 has 393,886 rows over the four moments, close to the average of 384,445 rows per block. The Python measurements of block 1 were repeated eight times.
one block
Block 1, four departure moments, 393,886 rows, in seconds; the fastest and the median run:
| step | Python, one process (8 runs) | GeoDMS, inside the output step (3 runs) |
|---|---|---|
| read the four files (53 MB, from the file cache) | 0.02 / 0.03 | – |
| parse the CSV text | 0.41 / 0.57 | – |
| the counts themselves | 0.32 / 0.46 | 0.053 / 0.054 |
| total | 0.73 / 1.03 | 0.053 / 0.054 |
The Python times of the counts are the full script minus parsing alone. The slowest Python run took 1.71 s, at a machine load of 96%.
The GeoDMS times per departure moment were 7 to 14 ms, of which unique over the 98 thousand keys took 3.5 to 7 ms, and rlookup, pcount and each any 0.3 to 3 ms. The union over the moments (128 thousand pair rows into 32 thousand pairs) took 8 to 12 ms, except in one run, in which a single unique took 254 ms and the block 0.36 s in total; with another run competing for the machine, an operation can wait for memory or a thread.
For scale: the whole output of block 1, from reading the chain stores to writing the four CSV files, took 8.5 to 11 minutes, or 573 to 893 seconds summed over the operations. The counts are less than 0.01% of that (0.06% in the slow run).
the whole morning
All 450 blocks, 1,800 files, 23.3 GB, 173,000,050 rows:
| Python, afterwards | GeoDMS, in the output step | |
|---|---|---|
| input | 1,800 CSV files, read and parsed | none: the results are in memory |
| wall time | 65 s, with 8 processes | not separately visible: the output step without the counts took 2 h 36 min |
| CPU time | 491 s | about 25 s (450 blocks x 54 ms), so at most 25 s of the 2 h 36 min |
| extra memory | at most 61 MB per process | about 2 MB per departure moment of a block |
where the difference comes from
- Reading and parsing. Of each 31-column row, Python makes new string objects for every field and uses five. That was about 35 to 55% of its time. The GeoDMS reads nothing: the counts use arrays the output step already holds.
- Types. Python identifies an OD pair by a tuple of two strings, which it allocates, hashes and stores per row in up to four sets. In the GeoDMS a pair is one element of the combined domain
OD_code, stored as one 8-byte integer, anduniqueover 98 thousand integers takes a few milliseconds. - Interpretation. The Python loop executes a dozen interpreted operations per row; a GeoDMS operator is one compiled loop over a typed array. For the counting itself, the GeoDMS was 6 to 9 times faster on the same rows; including the parsing that it does not need, 14 to 19 times.
- Parallelism at a different grain. Python runs independent blocks in separate processes; the GeoDMS runs operations, and segments of large operations, in threads. The output step keeps all cores busy anyway, so the counts add no wall time worth mentioning.
what to take from this
- Count where the data is. When the data is in memory anyway, a summary in the model costs a few dozen lines of configuration and, here, well under 1% of the output step. It is made in every run, it is always consistent with the output, and it is traceable like any other item.
- A short script is right for data that already exists. One minute on 8 cores for 173 million rows of text is fine for a one-off or retrospective check. Stream the rows, keep the memory per worker small, and run independent files in parallel. When you have to read the same output often, a binary column format such as parquet (Module 2D, GeoDMS and Python) or a native GeoDMS format (Module 2C, GeoDMS own data formats) avoids most of the parsing.
- Two implementations are also a test. The Python script and the GeoDMS code were written separately from the same definitions. For block 1 they gave identical counts, which checks both.
try it yourself!
-
Assemble the example of this page in one file: the skeleton, with the two classifications in
Classificationsand the threeTellingencontainers at the places marked see above and see below. Open it in the GeoDMS GUI and update/NetworkSetup/ConfigurationPerBlock/Generate_Output/OUTPUT_Generate_PublicTransport_fullOD_long_CSVFiles, or run that item withGeoDmsRun. The folder%LocalDataProjDir%/Outputthen hasod_tellingen.csvandunieke_od_tellingen.csv, together:soort;07h00m;07h15m;totaal alle;2000;2200;4200 OV-keten;1332;1466;2798 WW;224;246;470 CC;222;244;466 AA;222;244;466 groep;07h00m;07h15m;samen alle;584;584;1060 OV-keten;584;584;1060 WW;212;230;416 CC;208;226;420 AA;204;224;408 zonder auto;484;484;860 zonder auto en fiets;424;424;768Then change the test data, for example
NeedsCar := i % 2 == 0, predict which numbers change, and run it again. - Make a large table yourself. In Python, write a CSV file with 10 million rows and two columns,
organddest, filled with names such asorg123anddest4567(draw the numbers at random below 10,000). Count the unique(org, dest)pairs with a set of tuples of strings, and time it withtime.perf_counter. Then turn the names into integers first (adictper column) and count the unique keysorg * 10000 + dest. How much of the time went into the tuples of strings, and how much into reading the file? - In the GeoDMS, read the same file (Module 2A). Make domain units of the unique origins and destinations with
unique, relate the rows to them withrlookup, combine the two relations into one withcombine_unit_uint64andcombine_data, and count the elements ofuniqueon it. Compare the count with your Python result, and look up the time of each step in View > Calculation times (Deepening II). - Add a condition, for example the destination number is even, as a
boolattribute, and count withany(condition, relation)the pairs that meet it at least once. Write both counts to a text file with aparameter<string>andStorageType = "str". - Think about a summary in your own model: is the data it needs in memory during a run anyway? Then which would you choose, a few lines in the model or a script afterwards?
Go to previous deepening: Deepening II, Performance and memory