Dimensions and lookups#
A dimension is an axis of the model, such as snapshot or generator.
Declarations are indexed by it, and sum reduces along it.
A lookup is a named map out of a dimension: one value for each of its members. A generator's bus is a lookup, and so is a snapshot's period.
dimensions#
Every dimension named anywhere in the file is declared here.
| Field | ||
|---|---|---|
dtype |
float, int, str, datetime |
default str |
description |
free text, never parsed | default null |
A declaration says that the axis exists and what type its labels have. It never lists the labels. The generators, buses and snapshots arrive with the data, as a table, not as a list somebody keeps in step by hand.
Where the members come from#
The engine that binds the data follows three rules, and every engine follows the same three. So two engines given the same file and the same tables build the same model.
- The members come from the key named after the dimension. An engine reads
generatorfrom thegeneratortable, and from nowhere else. It readsp_maxfor its values, never for its list of generators, and it does not treatgen_busas the list either. If a declaration usesgeneratorand nogeneratortable arrives, the engine raises an error that namesgenerator. It does not build an empty axis, because an empty axis would silently delete every row indexed by it. A declared dimension that no declaration uses needs no table. - The members keep the order the table gives them. The engine does not
sort them, whether they are strings, integers or dates.
shift,sum_backandposition()all count along this order, so an engine that sortedsnapshotwould giveshift(p, over=snapshot, offset=1)a different meaning. To get a particular order, write the table in that order. - A table has each coordinate at most once. Two rows for
snapshot == 3is an error that names3. The engine does not keep the last, keep the first, or add them. At most once, not exactly once: a coordinate with no row is absence, and absence is how a model masks. A lookup's table obeys the same rule.
Every dimension has one list of members, and every parameter is lined up against
it when the data binds. So if load has 8760 snapshots and price has 8759, the
engine raises an error rather than build a model with one snapshot dropped.
lookups#
A lookup is how the network's wiring stays in the data. Which bus each generator sits on, which two buses each line joins, which period a snapshot falls in: each is a lookup table, and the file holds no adjacency matrix.
Declare each lookup under its own name. over: names the dimension whose
members carry the value, and into: names the dimension the values are labels
of. sum(by=) and at(by=) land terms on that target
dimension:
dimensions:
bus: { dtype: str }
generator: { dtype: str }
line: { dtype: str }
snapshot: { dtype: int }
period: { dtype: int }
lookups:
gen_bus: { over: generator, into: bus }
line_from: { over: line, into: bus } # two lookups onto one dimension
line_to: { over: line, into: bus }
period_of: { over: snapshot, into: period }
| Field | ||
|---|---|---|
over |
required — the dimension whose members carry the map | |
into |
required — the dimension its values are labels of, other than over |
|
description |
free text, never parsed | default null |
The target must be a declared dimension, and it must differ from over. The
values are checked against it when the data binds, which is the check that makes
sum(by=) safe.
That check is also why a label set the model only ever selects on is declared
as a dimension all the same. Nothing above is indexed by period; a declaration
selects on it with where: "period_of == 1"
(where strings).
A partial lookup is legal. A label the map leaves out belongs to no group, so a
generator can sit on no bus and a line can have one open end. sum(by=) places
such a label's terms nowhere. A value that names no label of the target is an
error.
Several lookups may group at once. sum(x, by=[gen_bus, gen_tech]) groups
through both maps in one reduction and lands on bus and technology. Every
lookup in the list must be over: the same dimension, and each must target a
different one. A member that either map leaves out belongs to no group.
Every lookup name joins the flat namespace, so a lookup may not shadow a
dimension, and that includes its own target. The map from generator onto bus
is called gen_bus, never a second bus.
How the map is supplied#
The map is a source key like any other, under the lookup's own name. It carries
two columns, each named after the dimension it holds: the over dimension, and
the target:
sources = {
'generator': ['g1', 'g2', 'g3'],
'gen_bus': pl.DataFrame({'generator': ['g1', 'g2'], 'bus': ['north', 'south']}),
}
A partial map is exactly the rows it has: g3 appears in no row, so g3 sits
on no bus. A null in the value column is refused, because a missing row already
says the same thing. A key that matches no label of over is an error rather
than a new member.
Values are never inferred from the parameters that use the target. If they were, a mistyped label would extend the label set instead of being rejected.
The map touches no table but its own, so you can add a lookup to a model the way
you add a parameter. The index of the over dimension may carry other columns,
but a column named after the lookup is refused rather than read.
Dimension, lookup or parameter?#
Every column of data is one of the three. What decides which is what the math does with the column, not what the column holds:
| The column… | is declared as | because |
|---|---|---|
| is an axis: something is indexed by it, or an aggregation lands terms on it | a dimension |
its members are the coordinate set every table over it is reindexed onto |
| has one value per member of a dimension and points at another — a generator's bus | a lookup into that dimension |
it is a map that sum(by=) and at(by=) walk, and its values are checked against the target |
| is a label set the model only selects on or counts within — a period, a season, a zone | a dimension, and a lookup into it |
the membership check is worth one line and one member list |
| scales terms — a coefficient, a bound, an offset | a parameter (float or int) |
arithmetic is over numbers (dtype) |
| is a per-row attribute the math only selects on — a fuel, a constraint's sense | a str parameter |
it names rows rather than scaling them, and no set is declared to check its values against |
| is a mask | a bool parameter |
a bare name in a where is its own answer |
Two rules follow from the table. If b has one value per a, then b is a
lookup over a, and not a dimension: a foreach product over two
dimensions that depend on each other, cut back with a mask, is the shape that
lookups replaces.
And everything under dimensions: is an axis. A dimension is never legal where
a value belongs, because it is a coordinate space and not data. To use a
dimension's coordinates as data, declare a parameter over it.
python -m math_spec check advises on a declared dimension that nothing is
indexed by, nothing aggregates into and no lookup targets
(errors).