Ragged trajectory¶
This chapter is the on-disk encoding of a trajectory section on the Zarr V3 binding. The logical model stays a sequence of frames; the layout below is how that sequence is stored so a writer appends and a reader resolves a frame without rewriting history. Groups sit on the root; chunk and shard extents, codecs and the commit protocol are Chunking and packing.
The layout is append-first (a new frame extends arrays) and sparse
(a block records an update only when it changes). A ragged run — a block
whose row count differs per frame — is the same layout: non-uniform
diff(offset).
One group per section¶
Dense per-step data (step, time, meta/<key>) has one row per frame;
the cell and each block carry their own sparse index of the ordinals at
which they changed. File count and open cost follow the number of sections,
and nstep stays off that axis.
There is one encoder and one decoder.
Layout¶
trajectory
+-- sequence_schema
+-- nstep the commit marker
+-- (step_progression: {start, stride})
+-- (time_progression: {start, stride})
\-- (step: i64[nstep]) only when the numbering is not a progression
\-- (time: f64[nstep]) only when supplied, and not a progression
\-- (meta)
| \-- <key>: <dtype>[nstep][...]
| +-- meta_dtype
\-- (box)
| +-- (cell_defined: bool[])
| +-- (vectors, origin, boundary) a fixed cell from ordinal 0, as attributes
| \-- (step_index: u64[n_updates]) the arrays, once the cell changes
| \-- (vectors: f64[n_updates][3][3])
| \-- (origin: f64[n_updates][3])
| \-- (boundary: bool[n_updates][3])
\-- <block>
+-- (structural_shape)
+-- (uniform_rows)
+-- (dense_updates)
\-- (step_index: u64[n_updates]) absent while the block is regular
\-- (offset: u64[n_updates+1]) absent while the block is regular
\-- <column>: <dtype>[total_rows][...]
\-- (_validity) only for columns declared nullable
\-- (<column>: bool[total_rows])
Every array under trajectory/ is sharded; extents and codecs are
Chunking and packing. The layout is built so that the common
run — a fixed number of atoms moving every frame, a topology written once, a
fixed cell, frames dumped every k steps — costs one array per column and
nothing else: no index arrays, no step array, no box/ arrays.
The commit marker: nstep¶
The trajectory group's attribute nstep is the number of frames fully on
disk. It is written last: a writer lands every array of a commit before
replacing the group metadata (atomically), so a crash between two writes of
one commit costs the uncommitted frames and nothing else. A reader takes
nstep from that attribute (a non-negative JSON integer); every array is
read to its logical length and MAY be longer (a tail the writer had not
yet committed) but MUST NOT be shorter. A store from a writer that kept
no marker attribute is read to len(step) (zero frames when there is no
step array either). The progression attributes below are only ever written
in the same metadata update as nstep, so a store that carries
step_progression or time_progression without nstep is malformed
and a reader refuses it. The full protocol is
Chunking and packing.
step and time¶
While the step numbers are an arithmetic progression — the dump every k
steps that every MD run produces — they are recorded as the attribute
step_progression: {"start": s, "stride": k} and no step array
exists. stride is step[1] − step[0] and appears from the second frame
on (a one-frame sequence carries {"start": s}; a zero-frame one carries no
attribute). The first frame whose step number breaks the progression
materializes the step array, backfilled with the whole history, and the
attribute is removed. time follows the same rule with time_progression.
The arithmetic is normative, because a progression is only as exact as the formula that expands it:
step[i] = start + i × stride, exact integer arithmetic;startandstrideare JSON integers.time[i] = start + f64(i) × stride, evaluated in IEEE 754 binary64 —iconverted to binary64, one rounded multiplication, one rounded addition, no fused multiply-add — withstride = time[1] − time[0](one rounded subtraction). A writer keeps a time series as a progression only while every value equals that expression bit for bit; an accumulated time usually does not, and goes to the array. A non-finite time cannot be spelled in JSON, so a series holding one is always an array.
A reader resolves frame i from the attribute when present, else from the
array.
time exists only when the run supplied times. It is all-or-nothing: a
sequence that starts without times cannot gain them later, and one that
starts with them cannot drop them. A writer MUST refuse either.
step holds the producer's iteration counter (strictly increasing; a
repeated or smaller step number is refused at append). Every step_index
below holds frame ordinals 0 .. nstep-1, not those counters, and is
strictly ascending: a reader refuses one that repeats or decreases. See
Trajectory.
Per-step metadata¶
Frame metadata that varies per step lands as one typed array per key under
trajectory/meta/, of shape [nstep] for a scalar or [nstep][3|6|9] for a
fixed vector. Each array MUST carry the attribute meta_dtype, whose
value is the exact tag of the values it holds; a reader refuses an array
without one, and an array whose stored dtype or trailing shape is not the
one its tag names.
The tag set is closed — exactly these sixteen:
| Tag | Stored as, per step | Value |
|---|---|---|
bool |
bool |
a boolean |
i32, i64 |
i32, i64 |
a signed integer |
u32, u64 |
u32, u64 |
an unsigned integer |
f64 |
f64 |
a binary64 real |
string |
string |
a UTF-8 string |
json |
string |
one UTF-8 JSON document |
bool3 |
bool[3] |
three booleans |
i32x3, i64x3 |
i32[3], i64[3] |
three signed integers |
u32x3, u64x3 |
u32[3], u64[3] |
three unsigned integers |
f64x3, f64x6, f64x9 |
f64[3], f64[6], f64[9] |
a vector, a Voigt tensor, a 3×3 matrix (row-major) |
There is no f32 tag, no complex tag, and no other width: a value that
needs one belongs in a block column. json is a meta tag only: a block
column never carries it.
A value is held exactly to its tag. A writer coerces what a producer
hands it only where that is exact — an integer given for an f64 key is
that real — and refuses everything else: a real for an integer tag, an
integer out of the tag's range, a boolean for a number, a vector of the
wrong width, a json value that is not finite JSON. The
standard per-step scalar keys (pe, ke, etotal, temp, press,
volume, all f64) are Standardized identifiers.
A meta key is declared once, when the sequence is created, as
{dtype, fill?}. A frame that omits a declared key is an error unless
that key was declared with an explicit fill value. There is no implicit
NaN. The fill is materialized into the array at the omitting step and
the declaration — tag and fill — is recorded in the pinned
sequence_schema. The array itself does not
mark which steps were filled.
A fill is a value of its tag, held to the same rules, and written in the
typed JSON form ("NaN" for a NaN
f64, a decimal string for an integer beyond ±2⁵³). fill absent
declares no fill; fill: null declares the JSON document null and is
valid only for a json key.
The cell¶
box/ is one section, updated whenever the cell changes. Update j holds
the cell matrix (vectors[j], lattice vectors as columns), its origin
(origin[j]) and its per-axis periodic flags (boundary[j]), at frame
ordinal step_index[j]. As on a frame, origin
and boundary are arrays, optional with the normative defaults (zero
origin, all-periodic); the reference writer omits each when every update
holds the default. A trivial box/step_index — exactly one update at frame
ordinal 0 — may also be omitted; absence reads back as [0].
A fixed cell costs no array. A cell present from ordinal 0 that never
changes is recorded as the box/ group's own attributes — vectors (a 3×3
nested list, lattice vectors as columns), and origin / boundary when off
their defaults — with no arrays at all. The first change migrates the cell
into the arrays above (update 0 is the attribute cell, backfilled) and
removes the vectors / origin / boundary attributes. A reader that finds
no box/vectors array reads the attributes: one update at ordinal 0.
The group MAY carry the boolean attribute cell_defined. Absent means
true. A writer emits it only to record false. It is the section's one
flag, covering every update — an update carries no flag of its own, so the
two cannot disagree — and an undefined cell's updates follow the
undefined-cell rules: identity vectors (ignored
by a reader), all-false boundary.
The cell carries forward like a block: once a run writes a cell, every later
frame resolves to the most recent update. A frame that drops its cell
mid-run reads back carrying the previous cell; box/ has no absence marker.
Per-block sparse updates (CSR)¶
Each block of the frames is one group under trajectory/, holding the rows
that block contributed across the whole run, in Compressed Sparse Row form:
| Element | Type | Meaning |
|---|---|---|
B/step_index |
u64[n_updates] |
ascending frame ordinals at which B was updated |
B/offset |
u64[n_updates + 1] |
CSR prefix sum of row counts: offset[0] == 0; update j owns rows offset[j] … offset[j+1] of every column |
B/<column> |
dtype[total_rows][…trailing] |
columns appended update after update; total_rows == offset[n_updates] |
Row counts are never stored; they are diff(offset), so the two cannot
disagree. A conforming reader MUST compute the row count as
offset[j+1] − offset[j] with checked subtraction and MUST reject a
non-monotonic offset.
A block that declares a structural shape
mirrors it onto its group as the attribute structural_shape. Such a block
has a fixed row count: every update holds exactly the product of the
shape, and a writer refuses an update that does not.
Nullable columns¶
A column the declaration pins nullable carries its
validity mask as B/_validity/<column>:
bool[total_rows], one flag per row of the section, cut by the same
offset as the values, so update j's flags are rows
offset[j] … offset[j+1]. An update that presents the column without a mask
lands all-true flags. A frame that masks a column the declaration pins
non-nullable is refused: its values would land without the mask. Columns
the declaration does not pin nullable have no mask array.
A regular block costs no index. A block is regular while every update
so far has the same row count N > 0 and sits at its own ordinal
(step_index[j] == j). While it is regular a writer writes no
step_index and no offset and instead marks the block group with two
elision markers:
| Attribute | Meaning |
|---|---|
uniform_rows: N |
every update so far has exactly N > 0 rows |
dense_updates: true |
every update so far sits at its own ordinal |
The markers are a pair and they stand in place of the index — never beside it:
- Update
jis at ordinaljand owns rowsj·N … (j+1)·N. The number of updates ismin(⌊len(c) / N⌋, nstep)over the shortest columncof the block (floor division: a partial trailing update is an uncommitted tail). - The first update that breaks either rule — a different row count (a ragged run, a zero-row update), or an update at an ordinal other than its own (a topology that changed at frame 5) — materializes both arrays, backfilled with the regular history, and the same commit removes both markers for good. A committed store never carries a marker beside an index array; a reader that finds both (a crash between the two writes of that commit) uses the arrays, which are authoritative.
- A block with no columns has no length to count its updates by, so it always keeps its index.
- One marker without the other,
uniform_rowsthat is not a positive integer, or markers on a block with no columns is malformed, and a reader refuses it.
A constant topology (one update at ordinal 0) and coordinates that move
every frame are both regular; only a block that sometimes changes, or
changes size, pays for an index.
A declared block with no updates is a group with zero-length columns and neither markers nor an index (a writer may create every declared block's group when the sequence is created). It is absent at every ordinal; a reader MUST NOT refuse it. Rows its columns hold beyond the committed frames are an uncommitted tail, as everywhere else.
The three states of a block¶
At any frame ordinal i, a declared block B is in exactly one of three
states, decided by its most recent update at or before i:
| State | On disk | What the frame carries |
|---|---|---|
| present | the most recent update has rows | B with those rows |
| empty | the most recent update has zero rows (offset[j+1] == offset[j]) |
B with its declared columns and no rows — it does appear in the frame |
| absent | no update at or before i |
no B at all |
The rules that follow from this:
- A frame that omits a declared block writes no update. The block carries forward: at that ordinal it is whatever the previous ordinal made it. Omission is never written as a zero-row update.
- A frame that presents a zero-row block writes a zero-row update. From that ordinal on the block is present and empty, until its next update.
- A block is absent only before its first update. Once present it is never absent again: there are no tombstones. To clear a block, write a zero-row update.
- A writer compares a presented block with its previous update bitwise,
after rounding every column that declares a
precision; an identical presentation writes
no update. A repeated
NaNtherefore counts as unchanged, and so does a change smaller than half the quantum. - A reader resolving a zero-row update hands back a block with the declared columns and zero rows, and that block appears in the frame.
The conformance suite distinguishes the three states.
Raggedness¶
"Ragged" means the number of rows in a block is not constant across frames. Growth, grand-canonical insertion, reactive topology, and variable-size selection sets are all the same encoding: each update owns a (possibly different) slice of the concatenated columns.
A frame may omit a whole block. Every column of a block it presents is present: all columns of a block share one CSR row range.
System and trajectory side by side¶
When system/<block> and trajectory/<block> coexist, the trajectory
block is the per-step state of the rows system defines, aligned 1:1 by
row order: every update of trajectory/<block> holds exactly as many rows
as system/<block>. When both sides carry an id column the values are
equal row for row, so a reader may also join on id. A ragged trajectory
block MUST NOT share a name with a system block. See
Trajectory.
Aligned blocks¶
A block may declare aligned_with: "<target>" in its sequence_schema
entry. Its rows are then the target's rows, one for one, at every frame:
a sparse companion of a block that changes more often — atom types beside
coordinates under proton hopping, species under grand-canonical insertion.
- The target is a declared block of the same trajectory, not the aligned
block itself, and not itself aligned. The aligned block declares no
structural_shape, and no column name appears in both. - At every frame ordinal where the aligned block is present or empty, the target is present or empty with the same row count, both resolved by carry-forward. A frame whose target update changes the row count therefore restates the aligned block; a frame that keeps the count may let it carry forward.
- The aligned block may be absent while its target is present, never the reverse.
- Rows correspond by position. A producer that reorders or replaces rows
without changing their count restates the aligned block. (The two blocks
share no column, so an
idlives in one of them.) - A writer refuses a frame that breaks the rule; a reader refuses a store that does.
- A reader hands back two blocks. It MAY also offer them joined; the column sets are disjoint, so the join is the union of columns.
- An aligned block never shares a name with a
systemblock. Its target may: the target's row count is then thesystemblock's, and so is the aligned block's.
The pinned declaration¶
The set of blocks, columns, dtypes, trailing shapes and per-step meta keys
a trajectory may carry is declared when the trajectory is created and
fixed for its lifetime. A later frame MAY present a subset of the
declaration, but MUST NOT present a block, column, or meta key outside
it. A run that must record a new section that was not declared needs a new
trajectory.
To record a section that appears only partway through a run (per-frame bonds in a growth run), a writer declares it up front — by declaring the block explicitly, or by including a representative frame carrying an even-empty block when deriving the declaration.
Reserved names are refused at declaration, not at the first write (see Reserved names).
The declaration is pinned as the trajectory/ group attribute
sequence_schema, published as
schema/binding/sequence-schema.schema.json:
blocks: a map of block name to{columns, structural_shape?, targets?, aligned_with?}, each column{dtype, trailing, nullable?, precision?}. Columndtypes are the closed record dtype set;trailingis the per-entity shape after the leading count axis;nullable: truedeclares a nullable column and is written only when true (absent meansfalse);precisionis the column's declared precision, the only place a trajectory states it, written only when declared;targetsis the block's row references;aligned_withmakes it an aligned block. Each of the last three is written only when set.meta: a map of key to{dtype, fill?},dtypedrawn from the per-step tag set above.
sequence_schema.blocks is the authoritative block list. A reader
resolves exactly the declared blocks: a declared block whose group is
missing is absent at every ordinal, and a child group of trajectory/ the
declaration does not name (and that is not one of the reserved names) is
not a block of the sequence — a reader does not resolve it into frames, and
preserves it as unknown content. Only a store that carries no
sequence_schema at all (a foreign writer's) is read by deriving the
declaration from its groups.
The attribute carries no version of its own; the record's
meta["molrec_version"] covers it.
Resolving a frame¶
To read block B at frame ordinal i:
- binary-search
B/step_indexfor the largest entry≤ i; call it updatej; - no entry
≤ imeansBis absent at framei— no block, not an empty one; - otherwise the frame's rows are
offset[j] … offset[j+1]of each column ofB. A zero-row range yields an empty block: present, with its declared columns and no rows.
With the elision markers there is no index to search: j = min(i,
n_updates − 1) with n_updates = min(⌊len(c) / N⌋, nstep) over the
shortest column c, and the rows are j·N … (j+1)·N; with zero updates the
block is absent.
The box resolves the same way over box/step_index, minus the row range:
the frame's cell is update j itself.
Resolution therefore costs a binary search over a section's changes, not over the run length.
Reserved names¶
Two small namespaces are owned by the layout:
| Namespace | Reserved names | Owner |
|---|---|---|
Children of trajectory/ |
step, time, meta, box |
the sequence itself |
| Children of a block group | offset, step_index, _validity |
that block's index and masks |
A block named step / time / meta / box, or a column named offset /
step_index / _validity, is rejected when the sequence is declared — not
at the first write.
Cost tracks change¶
Because every section carries its own index — or, while it is regular, none at all — the cost of a section follows how often it changes, not how long the run is:
| Section | Updates | On disk beyond its columns |
|---|---|---|
Constant topology (bonds that never changes) |
1 | nothing: regular, two marker attributes |
Coordinates (atoms, changing every step, fixed count) |
nstep |
nothing: regular, two marker attributes |
| Reactive / grand-canonical topology | one per change | step_index and offset, one entry per change |
Ragged run (atoms gaining rows) |
one per change | step_index and offset, one entry per change |
Fixed cell (box/ under NVT) |
1 | nothing: attributes on box/ (vectors, plus origin / boundary when off their defaults) |
Changing cell (box/ under NPT) |
one per change | step_index, vectors (+ origin / boundary when off their defaults) |
step dumped every k steps |
— | nothing: step_progression |
A section earns a new update only when its content differs bit for bit
from its previous update: a repeated NaN is unchanged, and a -0.0 after
a 0.0 is a change.
This is what makes hand-splitting an unchanging topology into system/
optional, and it is what lets a topology that changes sometimes be
expressed at all.
Worked example: a growth trajectory¶
A growth trajectory (growth.mrec/) whose atoms block gains rows over 19
frames, with only positions, an element string, and a molecule id; frames
dumped every 10 steps from step 0; a constant cell; the potential energy per
step.
growth.mrec
\-- meta
| +-- molrec_version: 1
\-- trajectory
+-- sequence_schema
+-- nstep: 19
+-- step_progression: {"start": 0, "stride": 10}
\-- meta
| \-- pe: f64[19]
| +-- meta_dtype: "f64"
\-- atoms
| \-- step_index: u64[19]
| \-- offset: u64[20]
| \-- element: string[total_rows]
| \-- mol_id: u64[total_rows]
| \-- x: f64[total_rows]
| \-- y: f64[total_rows]
| \-- z: f64[total_rows]
\-- box
+-- vectors: [[20, 0, 0], [0, 20, 0], [0, 0, 20]]
meta/ carries molrec_version: 1: writers always create it and stamp the
version. There is no step array (the numbering is a progression) and no
box/ array (the cell is fixed from ordinal 0, so it is three attributes —
here only vectors, because origin and boundary hold their defaults).
atoms is not regular — its row count changes — so it carries its index
and no markers. Suppose the first three frames have 2, 2, and 5 atoms. Then:
Frame 0 owns rows 0:2, frame 1 owns 2:4 (still 2 atoms — the count
did not grow, but positions changed, so there is still an update), frame 2
owns 4:9. Had this writer also declared bonds and presented the same
bonds on every frame, bonds would be regular — one update at ordinal 0,
uniform_rows and dense_updates: true on its group, no index — and a
frame that omitted them would still read them back.
The pinned trajectory attribute:
{
"sequence_schema": {
"blocks": {
"atoms": {
"columns": {
"element": { "dtype": "string", "trailing": [] },
"mol_id": { "dtype": "u64", "trailing": [] },
"x": { "dtype": "f64", "trailing": [] },
"y": { "dtype": "f64", "trailing": [] },
"z": { "dtype": "f64", "trailing": [] }
}
}
},
"meta": {
"pe": { "dtype": "f64" }
}
}
}
Frame i is read by resolving each block through the algorithm above:
atoms at frame i owns rows offset[j] … offset[j+1] for the update j
that step_index finds (here j == i, since atoms change every frame); the
step number is 0 + 10·i; the cell is the one attribute cell.