Layout by example¶
This chapter is a guided tour of how MolRec data sits on disk. It walks through small, concrete records — one water molecule, two short trajectories, a force field, a run, a collection — and shows each one as the tree of groups, arrays and attributes the reference codecs actually write.
The tour explains; it does not legislate. Every rule it mentions is stated in the chapter it links to, and where the two seem to differ, the linked chapter wins.
Every tree here is generated. scripts/layout_examples.py builds each
example from molrec's own models, writes it with molrec's own codecs (the
Zarr V3 binding, the LMDB binding) into a temporary directory, and reads what
landed on disk back out: each node's zarr.json, the leading values of each
array, the keys and value bytes of the LMDB file. python
scripts/layout_examples.py --write regenerates the blocks of this page, and
a test fails when they drift from what the codecs write.
How to read the trees¶
The trees use the notation of Objective:
<name>/ a group (a directory with a zarr.json)
+-- key: value an attribute of the group above, as JSON
\-- child/ a child group
\-- name: f64[3] = [1.0, …] an array: dtype, shape, its leading values
| zarr: chunk [3] · bytes → crc32c
| how the array is chunked and encoded
The zarr: line abbreviates the array's metadata. chunk [n] is the inner
chunk shape; shard [m], chunk [n] means the array uses the
sharding_indexed codec, m rows per shard file and n rows per inner
chunk, with the shard index at the start of the file. The codec pipeline
reads left to right as the writer applies it:
| Abbreviation | Codec |
|---|---|
bytes |
raw little-endian bytes |
vlen-utf8 |
variable-length UTF-8 strings |
shuffle(8) |
numcodecs.shuffle, element size 8 — a byte shuffle |
gzip(1), zstd(3) |
lossless compressors, with their level |
crc32c |
a checksum that closes every chunk |
An ellipsis … marks values or long attributes cut short for the page.
The 30-second picture¶
One record is one directory, x.mrec/, and that directory is a Zarr V3
root: a zarr.json at its top, and below it one group per section. Every
group and every array is a directory holding its own zarr.json; an array's
data sits beside it as chunk files (x/c/0, x/c/1, …). Packed for storage,
the same files become the entries of one x.mrec.zip
(below).
The sections come in a few kinds, and the kind decides the shape on disk:
x.mrec/
(zarr.json: the Zarr V3 root)
|
+-------------------+---------+----------+-----------------------+
| | | |
documents frame-shaped sequence arrays and runs
attributes only attributes + blocks blocks over time
| | | |
meta/ frame/ trajectory/ observables/
status/ system/ metrics/
method/ forcefield/
- A document is a JSON object stored as the attributes of an otherwise
empty group:
meta(identity, always present),status(lifecycle) andmethod(how a job ran). - A frame-shaped section is a group whose attributes are a document and
whose child groups are blocks — tables of columns sharing one row count.
frameis one snapshot,systemthe topology,forcefieldthe parameter tables. - The sequence
trajectoryholds blocks over time, written only when they change. observablespairs named result arrays with their metadata;metricsholds training curves as dense arrays plus a live text log.
Here is a record carrying every section at once, more than a real record would (a frame of water beside a hydrogen peroxide system): each root group with its kind, its attribute keys and its children.
tour.mrec/ the root group; its attributes: {}
\-- meta/ document attributes: molrec_version, creator, created_at
\-- system/ frame-shaped attributes: smiles, total_charge, _meta_types
| children: atoms/, bonds/, angles/, dihedrals/
\-- frame/ frame-shaped attributes: smiles, total_charge, units,
| _meta_types
| children: atoms/, bonds/, density/, box/
\-- trajectory/ sequence attributes: sequence_schema, step_progression,
| time_progression, nstep
| children: meta/, atoms/, box/
\-- forcefield/ frame-shaped attributes: name, units, special_bonds, styles
| children: angle.harmonic/, atom.full/,
| bond.harmonic/, dihedral.periodic/,
| pair.lj%2Fcut/
\-- observables/ array section children: meta/, dipole, msd
\-- method/ document attributes: type, description, engine
\-- status/ document attributes: state, stage, global_step, started_at,
| finished_at
\-- metrics/ catalog + series attributes: wal, series
children: series/, steps/, wall_time/,
metrics.jsonl
The rules: Overview for the section kinds, Root layout for how each lands on Zarr.
A single frame: one water molecule¶
A frame is one snapshot. This record holds one water molecule: its atoms,
its two bonds, a small density grid and a cubic box.
water.mrec/
\-- meta/
| +-- molrec_version: 1
\-- frame/
+-- smiles: "O"
+-- total_charge: 0
+-- units: {"length": "angstrom", "mass": "dalton"}
+-- _meta_types: {"smiles": "string", "total_charge": "i64", "units": "json"}
\-- atoms/
| +-- count: 3
| \-- b_factor: f64[3] = [11.8, 0.0, 0.0]
| | zarr: chunk [3] · bytes → crc32c
| \-- element: string[3] = ["O", "H", "H"]
| | zarr: chunk [3] · vlen-utf8 → gzip(1) → crc32c
| \-- mass: f64[3] = [15.999, 1.008, 1.008]
| | zarr: chunk [3] · bytes → crc32c
| \-- x: f64[3] = [0.0, 0.0, 0.0]
| | zarr: chunk [3] · bytes → shuffle(8) → zstd(3) → crc32c
| | +-- precision: 0.001
| \-- y: f64[3] = [0.0, 0.7568359375, -0.7568359375]
| | zarr: chunk [3] · bytes → shuffle(8) → zstd(3) → crc32c
| | +-- precision: 0.001
| \-- z: f64[3] = [0.1171875, -0.46875, -0.46875]
| | zarr: chunk [3] · bytes → shuffle(8) → zstd(3) → crc32c
| | +-- precision: 0.001
| \-- _validity/
| \-- b_factor: bool[3] = [true, false, false]
| zarr: chunk [3] · bytes → gzip(1) → crc32c
\-- bonds/
| +-- count: 2
| \-- atomi: u64[2] = [0, 0]
| | zarr: chunk [2] · bytes → gzip(1) → crc32c
| \-- atomj: u64[2] = [1, 2]
| | zarr: chunk [2] · bytes → gzip(1) → crc32c
| \-- bond_type: u64[2] = [1, 1]
| zarr: chunk [2] · bytes → gzip(1) → crc32c
\-- density/
| +-- count: 8
| +-- structural_shape: [2, 2, 2]
| \-- rho: f64[8] = [0.0313, 0.0449, 0.0388, 0.0113, 0.015, 0.0437, …]
| zarr: chunk [8] · bytes → crc32c
\-- box/
\-- vectors: f64[3][3] =
[[10.0, 0.0, 0.0], [0.0, 10.0, 0.0], [0.0, 0.0, 10.0]]
zarr: chunk [3, 3] · bytes → crc32c
Reading it top to bottom:
meta/comes first and carriesmolrec_version: 1. Every writer creates it and stamps the version, even when the producer gave no identity at all.- The attributes of
frame/are the frame's meta document.smiles,total_chargeandunitsare the producer's keys; there is nometachild group._meta_typesis the one reserved attribute: it tags each key with its type, sototal_chargereads back as ani64andunitsas a JSON document rather than whatever a JSON parser guesses. - A block is a group; its columns are its arrays.
atoms/carriescount: 3, and every column under it has 3 rows on its leading axis. Coordinates are three separate columnsx,y,z, never one packed array. - Relations point at rows.
bonds/atomiandbonds/atomjare 0-basedu64row indices intoatoms: the bonds are O–H1 and O–H2.bond_type1is a single bond. - A volumetric block declares its shape.
density/hascount: 8andstructural_shape: [2, 2, 2]; its columnrhois stored flat, and the attribute is what lets a reader reshape it into the grid. - A nullable column carries a mask. The structure resolves the oxygen
only, so
b_factorhas a value for row 0 alone. Its maskatoms/_validity/b_factoris[true, false, false]; the0.0stored under the hydrogen rows means nothing. Columns with no null rows (all the others) have no mask array at all. - A declared precision rounds before it encodes.
x,y,zcarry the attributeprecision: 0.001. The writer stored each value on the grid of the largest power of two not above it,q = 2⁻¹⁰: the input0.7572became0.7568359375(= 775/1024), withinq/2 ≈ 0.0005of what was given. Those arrays alone getshuffle(8) → zstd(3), because on that grid their low mantissa bits are zeros a compressor removes.massdeclares no precision: its bytes go to disk as they are. - The box is a reserved group, not a block. Only
vectorsis written (the lattice vectors are its columns);originandboundaryare omitted because they hold their defaults (zero, periodic on every axis).
Every array here is one chunk: a frame-path array is cut into 512 KiB pieces
along its leading axis, and is sharded only beyond four of them. One detail
that surprises people: the molecule lies in the yz plane, so x is all
zeros, and Zarr writes no chunk file for a chunk that holds only the fill
value. A reader gets the zeros back all the same
(the packed listing shows the missing file).
The rules: Frame-shaped group, Containers, Nullable columns, Declared precision, Simulation box, Standardized identifiers, and the extents and codecs in Chunking and packing.
A trajectory with fixed topology¶
Four frames of a vibrating hydrogen peroxide molecule, H–O–O–H. Nothing about
the molecule changes but its coordinates, so the topology goes in system,
written once, and the trajectory carries positions only. Frames are dumped
every 100 steps, 0.5 time units apart, in a fixed cubic cell, with the
potential energy and temperature of each frame.
vibration.mrec/
\-- meta/
| +-- molrec_version: 1
\-- system/
| +-- smiles: "OO"
| +-- total_charge: 0
| +-- _meta_types: {"smiles": "string", "total_charge": "i64"}
| \-- atoms/
| | +-- count: 4
| | \-- element: string[4] = ["H", "O", "O", "H"]
| | | zarr: chunk [4] · vlen-utf8 → gzip(1) → crc32c
| | \-- type: string[4] = ["H1", "O1", "O1", "H1"]
| | zarr: chunk [4] · vlen-utf8 → gzip(1) → crc32c
| \-- bonds/
| +-- count: 3
| \-- atomi: u64[3] = [0, 1, 2]
| | zarr: chunk [3] · bytes → gzip(1) → crc32c
| \-- atomj: u64[3] = [1, 2, 3]
| | zarr: chunk [3] · bytes → gzip(1) → crc32c
| \-- type: string[3] = ["O1-H1", "O1-O1", "O1-H1"]
| zarr: chunk [3] · vlen-utf8 → gzip(1) → crc32c
\-- trajectory/
+-- sequence_schema: {
| "blocks": {
| "atoms": {
| "columns": {
| "x": {"dtype": "f64", "trailing": []},
| "y": {"dtype": "f64", "trailing": []},
| "z": {"dtype": "f64", "trailing": []}
| }
| }
| },
| "meta": {"pe": {"dtype": "f64"}, "temp": {"dtype": "f64"}}
| }
+-- step_progression: {"start": 0, "stride": 100}
+-- time_progression: {"start": 0.0, "stride": 0.5}
+-- nstep: 4
\-- meta/
| \-- pe: f64[4] = [-12.4, -12.1476, -12.1272, -12.3577]
| | zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
| | +-- meta_dtype: "f64"
| \-- temp: f64[4] = [300.0, 302.5, 297.1, 300.0]
| zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
| +-- meta_dtype: "f64"
\-- atoms/
| +-- uniform_rows: 4
| +-- dense_updates: true
| \-- x: f64[16] = [0.8397, -0.0102, 0.0114, -0.8759, 0.8526, 0.0093, …]
| | zarr: shard [8388608], chunk [2048] · bytes → crc32c
| \-- y: f64[16] = [0.9072, 0.728, -0.7351, -0.8487, 0.8773, 0.7505, …]
| | zarr: shard [8388608], chunk [2048] · bytes → crc32c
| \-- z: f64[16] = [0.4465, -0.0105, 0.0149, 0.4201, 0.4144, -0.0041, …]
| zarr: shard [8388608], chunk [2048] · bytes → crc32c
\-- box/
+-- vectors: [[20.0, 0.0, 0.0], [0.0, 20.0, 0.0], [0.0, 0.0, 20.0]]
What to notice:
system/is frame-shaped, exactly likeframe/above: element, force-field type and bonds, written once. The trajectory'satomsrows line up 1:1 by position withsystem/atoms— every frame has 4 rows.trajectory/atomsis a regular block, so it has no index. Each column holds the 4 frames' rows one after the other, 16 rows in all. The group's attributesuniform_rows: 4anddense_updates: truesay that updatejsits at framejand owns rows4j … 4j+4, so neitherstep_indexnoroffsetis written.- The step and time numbers are formulas, not arrays.
step_progression: {"start": 0, "stride": 100}meansstep[i] = 100·i;time_progressionlikewise givestime[i] = 0.5·i. There is nosteportimearray. - A fixed cell costs no array.
box/holds the cell as its attributevectors. - Per-step scalars are one array per key.
meta/peandmeta/temphave one value per frame, and each states its type inmeta_dtype. sequence_schemais the pinned declaration. It lists every block, column, dtype and trailing shape the trajectory may ever carry, and every per-step key with its tag. It is fixed when the trajectory is created.nstep: 4is the commit marker. A writer lands every array first and replaces the group's metadata —nstepwith the progressions — last, so a crash costs the uncommitted frames and nothing else.
Every array under trajectory/ is sharded. A block column's inner chunk is
a whole number of frames of at least 16 KiB — here 2048 rows, 512 frames of
4 atoms — and a shard holds up to 4096 such chunks. Per-step and index
arrays use 1024-row chunks, 256 per shard. These are capacities: a shard file
holds its index and the chunks written so far, nothing for the chunks still
to come.
The rules: Trajectory, Ragged trajectory (the regular-block elision, step and time, the cell, per-step metadata, the pinned declaration, the commit marker), and Chunking and packing.
A trajectory whose topology changes¶
Five frames in which the chemistry moves. Two OH radicals drift; at frame 2 they recombine into H–O–O–H, which adds a bond and retypes the atoms; at frame 3 a grand-canonical move inserts an argon atom; at frame 4 the barostat resizes the cell. The producer handed the writer these frames:
| frame | step | blocks the producer presented |
|---|---|---|
| 0 | 0 | atoms (4), bonds (2: O–H, O–H), atom_types (4: HO, OH, OH, HO) |
| 1 | 10 | atoms (4) |
| 2 | 20 | atoms (4), bonds (3: the O–O bond forms), atom_types (4: H1, O1, O1, H1) |
| 3 | 25 | atoms (5: argon inserted), atom_types (5), insertions (1: row 4) |
| 4 | 30 | atoms (5), insertions (0 rows); cell 12 → 12.25 |
A block a frame does not present is not "missing": it carries forward from the previous frame.
reaction.mrec/
\-- meta/
| +-- molrec_version: 1
\-- trajectory/
+-- sequence_schema: {
| "blocks": {
| "atoms": {
| "columns": {
| "element": {"dtype": "string", "trailing": []},
| "x": {"dtype": "f64", "trailing": []},
| "y": {"dtype": "f64", "trailing": []},
| "z": {"dtype": "f64", "trailing": []}
| }
| },
| "bonds": {
| "columns": {
| "atomi": {"dtype": "u64", "trailing": []},
| "atomj": {"dtype": "u64", "trailing": []}
| }
| },
| "atom_types": {
| "columns": {"type": {"dtype": "string", "trailing": []}},
| "aligned_with": "atoms"
| },
| "insertions": {"columns": {"atomi": {"dtype": "u64", "trailing": []}}}
| },
| "meta": {"pe": {"dtype": "f64"}}
| }
+-- nstep: 5
\-- meta/
| \-- pe: f64[5] = [-4.1, -4.3, -12.2, -12.4, -12.3]
| zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
| +-- meta_dtype: "f64"
\-- step: i64[5] = [0, 10, 20, 25, 30]
| zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
\-- atoms/
| \-- step_index: u64[5] = [0, 1, 2, 3, 4]
| | zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
| \-- offset: u64[6] = [0, 4, 8, 12, 17, 22]
| | zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
| \-- element: string[22] = ["H", "O", "O", "H", "H", "O", …]
| | zarr: shard [8396800], chunk [2050] · vlen-utf8 → gzip(1) → crc32c
| \-- x: f64[22] = [0.8501, -0.0464, 0.0182, -0.8224, 0.8329, -0.0046, …]
| | zarr: shard [8396800], chunk [2050] · bytes → crc32c
| \-- y: f64[22] = [0.8844, 0.7426, -0.7219, -0.874, 0.9102, 0.7195, …]
| | zarr: shard [8396800], chunk [2050] · bytes → crc32c
| \-- z: f64[22] = [0.4208, -0.0425, 0.0166, 0.4113, 0.4104, -0.0103, …]
| zarr: shard [8396800], chunk [2050] · bytes → crc32c
\-- bonds/
| \-- step_index: u64[2] = [0, 2]
| | zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
| \-- offset: u64[3] = [0, 2, 5]
| | zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
| \-- atomi: u64[5] = [0, 2, 0, 1, 2]
| | zarr: shard [8392704], chunk [2049] · bytes → gzip(1) → crc32c
| \-- atomj: u64[5] = [1, 3, 1, 2, 3]
| zarr: shard [8392704], chunk [2049] · bytes → gzip(1) → crc32c
\-- atom_types/
| \-- step_index: u64[3] = [0, 2, 3]
| | zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
| \-- offset: u64[4] = [0, 4, 8, 13]
| | zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
| \-- type: string[13] = ["HO", "OH", "OH", "HO", "H1", "O1", …]
| zarr: shard [4198400], chunk [1025] · vlen-utf8 → gzip(1) → crc32c
\-- insertions/
| \-- step_index: u64[2] = [3, 4]
| | zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
| \-- offset: u64[3] = [0, 1, 1]
| | zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
| \-- atomi: u64[1] = [4]
| zarr: shard [8388608], chunk [2048] · bytes → gzip(1) → crc32c
\-- box/
\-- step_index: u64[2] = [0, 4]
| zarr: shard [262144], chunk [1024] · bytes → gzip(1) → crc32c
\-- vectors: f64[2][3][3] =
[[[12.0, 0.0, 0.0], [0.0, 12.0, 0.0], [0.0, 0.0, 12.0]], …]
zarr: shard [262144, 3, 3], chunk [1024, 3, 3] · bytes → gzip(1) → crc32c
Each block is its own sparse update series in compressed-sparse-row
(CSR) form: step_index lists the frame ordinals at which the block was
written, and update j owns rows offset[j] … offset[j+1] of every column.
atomschanges every frame and changes size at frame 3, so it is not regular: it carries the index,step_index = [0, 1, 2, 3, 4]andoffset = [0, 4, 8, 12, 17, 22]. Row counts are never stored; they are the differences ofoffset.bondswas written twice, at frames 0 and 2. Frames 1, 3 and 4 did not present it, so they cost nothing and read back the previous bonds.atom_typesis aligned withatoms(aligned_withinsequence_schema): its rows are the atoms' rows. It is written at frame 0, at frame 2 (retyped) and at frame 3 — restated becauseatomsgrew — and carries forward at frames 1 and 4, where the atom count did not change. Coordinates that move every frame do not drag the types along.insertionsshows absent versus empty. Before frame 3 it has no update, so frames 0–2 have noinsertionsblock at all. At frame 3 it has one row. Frame 4 presented it with zero rows, which is an update too:offset = [0, 1, 1], and from frame 4 on the block is present and empty. Once a block has appeared it is never absent again.- The step numbers are not a progression (an extra dump at step 25), so
they are the
steparray. - The cell changed, so
box/holds arrays:step_index = [0, 4]and onevectorsmatrix per change.
How a reader resolves frame i¶
To read block B at frame i, a reader binary-searches B/step_index for
the largest entry ≤ i — call its position j. No such entry means the
block is absent at frame i. Otherwise the frame's rows are
offset[j] … offset[j+1]; a zero-length range is an empty block, present
with its columns and no rows. A regular block has no index to search: its
update is j = min(i, n_updates − 1) and its rows are j·N … (j+1)·N. The
cell resolves the same way over box/step_index, without a row range.
This table is that algorithm run over the arrays printed above:
frame i |
step[i] |
atoms |
bonds |
atom_types |
insertions |
box |
|---|---|---|---|---|---|---|
| 0 | 0 | j=0 → 0:4 (4 rows) |
j=0 → 0:2 (2 rows) |
j=0 → 0:4 (4 rows) |
absent | j=0 (edge "12.0") |
| 1 | 10 | j=1 → 4:8 (4 rows) |
j=0 → 0:2 (2 rows) |
j=0 → 0:4 (4 rows) |
absent | j=0 (edge "12.0") |
| 2 | 20 | j=2 → 8:12 (4 rows) |
j=1 → 2:5 (3 rows) |
j=1 → 4:8 (4 rows) |
absent | j=0 (edge "12.0") |
| 3 | 25 | j=3 → 12:17 (5 rows) |
j=1 → 2:5 (3 rows) |
j=2 → 8:13 (5 rows) |
j=0 → 0:1 (1 row) |
j=0 (edge "12.0") |
| 4 | 30 | j=4 → 17:22 (5 rows) |
j=1 → 2:5 (3 rows) |
j=2 → 8:13 (5 rows) |
j=1 → 1:1 (empty) |
j=1 (edge "12.25") |
Take frame 3: in atoms/step_index = [0, 1, 2, 3, 4] the largest entry
≤ 3 is at position 3, so the frame's atoms are rows 12:17 of x, y,
z and element. In bonds/step_index = [0, 2] it is position 1: the
three bonds of rows 2:5, written at frame 2 and carried forward. The cost
of a lookup is a search over a block's changes, not over the run.
The rules: Per-block sparse updates, The three states of a block, Aligned blocks, Resolving a frame.
The force field¶
The parameters of the energy model live in their own frame-shaped section: the group's attributes are the force-field document, and each style is one table, stored as an ordinary block. This record carries the H–O–O–H system and a small force field for it.
peroxide.mrec/
\-- forcefield/
+-- name: "h2o2-demo"
+-- units: {"preset": "real"}
+-- special_bonds: {"lj": [0.0, 0.0, 0.5], "coul": [0.0, 0.0, 0.8333]}
+-- styles: [
| {"category": "atom", "style": "full"},
| {"category": "bond", "style": "harmonic"},
| {"category": "angle", "style": "harmonic"},
| {"category": "dihedral", "style": "periodic"},
| {
| "category": "pair",
| "style": "lj/cut",
| "params": {"cutoff": 10.0, "mixing": "geometric"}
| }
| ]
\-- angle.harmonic/
| +-- count: 1
| \-- name: string[1] = ["H1-O1-O1"]
| | zarr: chunk [1] · vlen-utf8 → gzip(1) → crc32c
| \-- itom: string[1] = ["H1"]
| | zarr: chunk [1] · vlen-utf8 → gzip(1) → crc32c
| \-- jtom: string[1] = ["O1"]
| | zarr: chunk [1] · vlen-utf8 → gzip(1) → crc32c
| \-- ktom: string[1] = ["O1"]
| | zarr: chunk [1] · vlen-utf8 → gzip(1) → crc32c
| \-- k: f64[1] = [100.0]
| | zarr: chunk [1] · bytes → crc32c
| \-- theta0: f64[1] = [1.7488]
| zarr: chunk [1] · bytes → crc32c
\-- atom.full/
| +-- count: 2
| \-- name: string[2] = ["H1", "O1"]
| | zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
| \-- charge: f64[2] = [0.41, -0.41]
| | zarr: chunk [2] · bytes → crc32c
| \-- element: string[2] = ["H", "O"]
| | zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
| \-- mass: f64[2] = [1.008, 15.999]
| | zarr: chunk [2] · bytes → crc32c
| \-- ptype: string[2] = ["A", "A"]
| zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
\-- bond.harmonic/
| +-- count: 2
| \-- name: string[2] = ["O1-O1", "O1-H1"]
| | zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
| \-- itom: string[2] = ["O1", "O1"]
| | zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
| \-- jtom: string[2] = ["O1", "H1"]
| | zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
| \-- k: f64[2] = [600.0, 1106.0]
| | zarr: chunk [2] · bytes → crc32c
| \-- r0: f64[2] = [1.475, 0.967]
| zarr: chunk [2] · bytes → crc32c
\-- dihedral.periodic/
| +-- count: 2
| \-- name: string[2] = ["H1-O1-O1-H1", "X-O1-O1-X"]
| | zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
| \-- itom: string[2] = ["H1", ""]
| | zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
| \-- jtom: string[2] = ["O1", "O1"]
| | zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
| \-- ktom: string[2] = ["O1", "O1"]
| | zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
| \-- ltom: string[2] = ["H1", ""]
| | zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
| \-- k1: f64[2] = [1.2, 0.5]
| | zarr: chunk [2] · bytes → crc32c
| \-- k2: f64[2] = [0.8, 0.0]
| | zarr: chunk [2] · bytes → crc32c
| \-- periodicity1: f64[2] = [1.0, 3.0]
| | zarr: chunk [2] · bytes → crc32c
| \-- periodicity2: f64[2] = [2.0, 0.0]
| | zarr: chunk [2] · bytes → crc32c
| \-- phase1: f64[2] = [0.0, 0.0]
| | zarr: chunk [2] · bytes → crc32c
| \-- phase2: f64[2] = [3.141592654, 0.0]
| | zarr: chunk [2] · bytes → crc32c
| \-- _validity/
| \-- k2: bool[2] = [true, false]
| | zarr: chunk [2] · bytes → gzip(1) → crc32c
| \-- periodicity2: bool[2] = [true, false]
| | zarr: chunk [2] · bytes → gzip(1) → crc32c
| \-- phase2: bool[2] = [true, false]
| zarr: chunk [2] · bytes → gzip(1) → crc32c
\-- pair.lj%2Fcut/
+-- count: 2
\-- name: string[2] = ["O1-O1", "H1-H1"]
| zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
\-- itom: string[2] = ["O1", "H1"]
| zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
\-- jtom: string[2] = ["O1", "H1"]
| zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
\-- epsilon: f64[2] = [0.17, 0.0]
| zarr: chunk [2] · bytes → crc32c
\-- sigma: f64[2] = [3.0, 1.0]
zarr: chunk [2] · bytes → crc32c
- The document is plain JSON in the group's attributes: the
name; theunitsevery number in the section is in (here therealpreset: ångström, kcal/mol, radian, e, dalton); the 1-2 / 1-3 / 1-4special_bondsweights; and the orderedstyleslist. A style may carry style-level parameters, aslj/cutdoes with itscutoffandmixingrule. - One block per style, named
<category>.<style>. Bytes of the style name outsideA–Z a–z 0–9 - _are percent-encoded, sopair/lj/cutlives atpair.lj%2Fcut. - A table has one row per type.
namelabels the row;itom…ltomname the atom types at its endpoints, as many as the category has (none foratom, two forbond, four fordihedral). Every other column is a parameter,f64orstring. - The empty string is a wildcard. The second torsion row,
X-O1-O1-X, has""at both ends: it matches any type there. - Parameters may be strings.
atom.full/ptypeis one. - A parameter a row lacks is null, not zero. The generic torsion has one
cosine term, the specific one two, so
k2,periodicity2andphase2carry masks[true, false]underdihedral.periodic/_validity/.
The system names its types; it never repeats the numbers:
peroxide.mrec/
\-- system/
+-- smiles: "OO"
+-- total_charge: 0
+-- _meta_types: {"smiles": "string", "total_charge": "i64"}
\-- atoms/
| +-- count: 4
| \-- element: string[4] = ["H", "O", "O", "H"]
| \-- type: string[4] = ["H1", "O1", "O1", "H1"]
\-- bonds/
| +-- count: 3
| \-- atomi: u64[3] = [0, 1, 2]
| \-- atomj: u64[3] = [1, 2, 3]
| \-- type: string[3] = ["O1-H1", "O1-O1", "O1-H1"]
\-- angles/
| +-- count: 2
| \-- atomi: u64[2] = [0, 3]
| \-- atomj: u64[2] = [1, 2]
| \-- atomk: u64[2] = [2, 1]
| \-- type: string[2] = ["H1-O1-O1", "H1-O1-O1"]
\-- dihedrals/
+-- count: 1
\-- atomi: u64[1] = [0]
\-- atomj: u64[1] = [1]
\-- atomk: u64[1] = [2]
\-- atoml: u64[1] = [3]
\-- type: string[1] = ["H1-O1-O1-H1"]
(Its arrays are stored exactly like the frame's above; their zarr: lines
are left out here.) Linking is by name: atoms.type names a row of the atom
table, and each relation block's type names a row of a table of its
category. Resolved against the tables read back from disk:
| system row | type |
table → row | parameters |
|---|---|---|---|
atoms[0] |
H1 |
atom.full row 0 |
charge 0.41, mass 1.008 |
atoms[1] |
O1 |
atom.full row 1 |
charge -0.41, mass 15.999 |
bonds[0] |
O1-H1 |
bond.harmonic row 1 |
k 1106.0, r0 0.967 |
bonds[1] |
O1-O1 |
bond.harmonic row 0 |
k 600.0, r0 1.475 |
angles[0] |
H1-O1-O1 |
angle.harmonic row 0 |
k 100.0, theta0 1.7488 |
dihedrals[0] |
H1-O1-O1-H1 |
dihedral.periodic row 0 |
k1 1.2, k2 0.8, periodicity1 1.0, periodicity2 2.0, phase1 0.0, phase2 3.141592654 |
Pair styles are resolved through the atom types instead: two O1 atoms take
the O1-O1 row of pair.lj%2Fcut, and a pair with no row of its own is
mixed from the two self rows by the style's mixing rule.
The rules: Force field — the document, style tables, endpoints, linking a system.
Observables and metrics¶
Observables are named scientific results. Each is a pair: a data array
observables/<name> of any shape, and its metadata document as the
attributes of observables/meta/<name>. One without the other is malformed.
results.mrec/
\-- observables/
\-- meta/
| \-- dipole/
| | +-- kind: "vector"
| | +-- description: "molecular dipole moment"
| | +-- time_dependent: false
| | +-- unit: "e*angstrom"
| \-- msd/
| +-- kind: "scalar"
| +-- description: "mean squared displacement"
| +-- time_dependent: true
| +-- unit: "angstrom**2"
| +-- sampling: "per_frame"
\-- dipole: f64[3] = [0.0, 0.0, 0.3929]
| zarr: chunk [3] · bytes → crc32c
\-- msd: f64[4] = [0.0, 0.0011, 0.0024, 0.0031]
zarr: chunk [4] · bytes → crc32c
kind says how to read the array (vector: one 3-vector; scalar: one
value per sample), and time_dependent: true says the leading axis of msd
is the trajectory's frames.
Metrics are run-local curves (training loss, validation error), kept
apart from scientific results. This run record is the
fixtures/run-minimal golden: meta, status and method documents, and
a metrics/ group holding both forms of the same three events.
fit.mrec/
\-- meta/
| +-- molrec_version: 1
| +-- creator: {"name": "molrec-fixtures", "version": "0.0.0"}
| +-- created_at: "2026-08-04T00:00:00+00:00"
\-- method/
| +-- type: "training"
| +-- description: "fit a potential to the H2O2 conformers"
| +-- engine: {"name": "molnex", "version": "0.0.0"}
\-- status/
| +-- state: "succeeded"
| +-- stage: "train"
| +-- global_step: 2
| +-- started_at: "2026-08-04T00:00:00+00:00"
| +-- finished_at: "2026-08-04T00:00:02+00:00"
\-- metrics/
+-- wal: {"lines": 3, "bytes": 233}
+-- series: {
| "train/loss": {
| "type": "scalar",
| "count": 2,
| "latest_step": 2,
| "latest_timestamp": "2026-08-04T00:00:02+00:00"
| },
| "eval/MAE": {
| "type": "scalar",
| "count": 1,
| "latest_step": 2,
| "latest_timestamp": "2026-08-04T00:00:02+00:00"
| }
| }
\-- series/
| \-- eval%2FMAE: f64[1] = [0.1]
| | zarr: chunk [1] · bytes → crc32c
| \-- train%2Floss: f64[2] = [0.5, 0.25]
| zarr: chunk [2] · bytes → crc32c
\-- steps/
| \-- eval%2FMAE: i64[1] = [2]
| | zarr: chunk [1] · bytes → gzip(1) → crc32c
| \-- train%2Floss: i64[2] = [1, 2]
| zarr: chunk [2] · bytes → gzip(1) → crc32c
\-- wall_time/
| \-- eval%2FMAE: string[1] = ["2026-08-04T00:00:02+00:00"]
| | zarr: chunk [1] · vlen-utf8 → gzip(1) → crc32c
| \-- train%2Floss: string[2] = ["2026-08-04T00:00:01+00:00", …]
| zarr: chunk [2] · vlen-utf8 → gzip(1) → crc32c
\-- metrics.jsonl (a plain file, not a Zarr node)
metrics/metrics.jsonl:
{"t":"scalar","k":"train/loss","s":1,"w":"2026-08-04T00:00:01+00:00","v":0.5}
{"t":"scalar","k":"train/loss","s":2,"w":"2026-08-04T00:00:02+00:00","v":0.25}
{"t":"scalar","k":"eval/MAE","s":2,"w":"2026-08-04T00:00:02+00:00","v":0.1}
- The live form is
metrics/metrics.jsonl, a plain UTF-8 file (not a Zarr node) with one JSON object per line, in compact keys:ttype,kseries key,sstep,wwall time,vvalue. A running job only ever appends to it. - The closed form is dense arrays, one per series key and per field:
series/holds the values,steps/the steps,wall_time/the times, aligned index for index. A key becomes an array name by percent-encoding:train/loss→train%2Floss. - The catalog is the attributes of
metrics/: a summary per series and the watermarkwal: {"lines": 3, "bytes": 233}— the dense arrays hold exactly the first 3 lines (233 bytes) of the log. Here that is the whole log; lines appended past the watermark would be newer, and a reader would add them.
The rules: Observables, Metrics, Metrics WAL, Status, Method.
A collection in LMDB¶
A training set is thousands of small records read in shuffled order. A collection stores them in one LMDB file, one key per record part, sharing one trajectory declaration and one force field. This one holds two H2O2 conformer records: record 0 with three frames (the third repeats the second's coordinates), record 1 with two.
conformers.mrec.lmdb (one file; keys in LMDB's byte order)
f ‖ u64be(0) key bytes 66 00 00 00 00 00 00 00 00
MRF1 · header 228 B · pad 4 B · payload 96 B
step 0 · meta {"pe": -12.4}
atoms[4]: buffers x f64 @0, y f64 @32, z f64 @64
f ‖ u64be(1) key bytes 66 00 00 00 00 00 00 00 01
MRF1 · header 228 B · pad 4 B · payload 96 B
step 1 · meta {"pe": -12.3}
atoms[4]: buffers x f64 @0, y f64 @32, z f64 @64
f ‖ u64be(2) key bytes 66 00 00 00 00 00 00 00 02
MRF1 · header 42 B · pad 6 B · payload 0 B
step 2 · meta {"pe": -12.2}
no blocks: nothing changed since the previous frame
f ‖ u64be(3) key bytes 66 00 00 00 00 00 00 00 03
MRF1 · header 228 B · pad 4 B · payload 96 B
step 0 · meta {"pe": -12.4}
atoms[4]: buffers x f64 @0, y f64 @32, z f64 @64
f ‖ u64be(4) key bytes 66 00 00 00 00 00 00 00 04
MRF1 · header 228 B · pad 4 B · payload 96 B
step 1 · meta {"pe": -12.3}
atoms[4]: buffers x f64 @0, y f64 @32, z f64 @64
ff key bytes 66 66
MRF1 · header 2495 B · pad 1 B · payload 232 B
meta keys name, units, special_bonds, styles
atom.full[2]: strings name, element, ptype; buffers mass f64 @0, charge f64 @16
bond.harmonic[2]: strings name, itom, jtom; buffers k f64 @32, r0 f64 @48
angle.harmonic[1]: strings name, itom, jtom, ktom; buffers k f64 @64, theta0 f64 @72
dihedral.periodic[2]: strings name, itom, jtom, ktom, ltom; buffers k1 f64 @80,
periodicity1 f64 @96, phase1 f64 @112, k2 f64 @128 (validity @144), periodicity2
f64 @152 (validity @168), phase2 f64 @176 (validity @192)
pair.lj%2Fcut[2]: strings name, itom, jtom; buffers epsilon f64 @200, sigma f64 @216
index key bytes 69 6e 64 65 78
MRF1 · header 353 B · pad 7 B · payload 56 B
records[2]: strings smiles; buffers first_frame u64 @0, n_frames u64 @16, n_atoms
u64 @32, has_trajectory bool @48
meta key bytes 6d 65 74 61
{
"layout": "mrec-lmdb",
"layout_version": 1,
"collection": {
"molrec_version": 1,
"units": {"length": "angstrom", "energy": "kcal/mol"}
},
"sequence_schema": {
"blocks": {
"atoms": {
"columns": {
"x": {"dtype": "f64", "trailing": [], "precision": 0.001},
"y": {"dtype": "f64", "trailing": [], "precision": 0.001},
"z": {"dtype": "f64", "trailing": [], "precision": 0.001}
}
}
},
"meta": {"pe": {"dtype": "f64"}}
},
"n_records": 2,
"n_frames": 5
}
r ‖ u64be(0) key bytes 72 00 00 00 00 00 00 00 00
{"record_id": "h2o2-0"}
r ‖ u64be(1) key bytes 72 00 00 00 00 00 00 00 01
{"record_id": "h2o2-1"}
s ‖ u64be(0) key bytes 73 00 00 00 00 00 00 00 00
MRF1 · header 523 B · pad 5 B · payload 48 B
meta {"smiles": "OO", "total_charge": 0} · meta_types {"smiles": "string",
"total_charge": "i64"}
atoms[4]: strings element, type
bonds[3]: strings type; buffers atomi u64 @0, atomj u64 @24
s ‖ u64be(1) key bytes 73 00 00 00 00 00 00 00 01
MRF1 · header 523 B · pad 5 B · payload 48 B
meta {"smiles": "OO", "total_charge": 0} · meta_types {"smiles": "string",
"total_charge": "i64"}
atoms[4]: strings element, type
bonds[3]: strings type; buffers atomi u64 @0, atomj u64 @24
- Keys are bytes, numbers are big-endian.
r,sandfkeys are one prefix letter followed by a big-endianu64, so LMDB's byte order is numeric order and one record's frames are one contiguous cursor range. metais JSON and is written last, in the transaction that commits the collection. It carries the layout tag and version, the collection document (withmolrec_versionand theunitsevery record's numbers are in), the onesequence_schemaevery record uses, and the counts.indexis one blockrecordswith a row per record: the binding's ownfirst_frame,n_frames,n_atoms,has_trajectory, plus the collection's columns (smiles). Record 1's frames start at global ordinal 3.r‖u64be(r)is recordr's meta document;s‖u64be(r)its system, which types its own meta (meta_types).f‖u64be(j)is the trajectory update at global frame ordinalj, holding only the blocks that changed there. Frame 2 changed nothing but its energy, so its value has no blocks at all; record 1's first frame (f‖u64be(3)) restatesatoms, because carry-forward never crosses a record boundary.ffis the collection's force field: the document as the frame's meta, one block per style table.
Every value except meta and r is frame bytes: a small fixed header,
a JSON description, then the numeric columns as raw buffers that a reader
can view in place. Here is f‖u64be(1), byte range by byte range:
key f ‖ u64be(1) = 66 00 00 00 00 00 00 00 01; value 336 B
[0, 4) magic 'MRF1'
[4, 8) header_length 228 (u32, little-endian)
[8, 236) header UTF-8 JSON:
{
"blocks": {
"atoms": {
"count": 4,
"structural_shape": null,
"columns": {
"x": {"dtype": "f64", "shape": [4], "offset": 0},
"y": {"dtype": "f64", "shape": [4], "offset": 32},
"z": {"dtype": "f64", "shape": [4], "offset": 64}
}
}
},
"meta": {"pe": -12.3},
"step": 1
}
[236, 240) padding 4 zero bytes
[240, 272) atoms/x f64 = [0.8212890625, -0.0517578125, -0.0078125, -0.8369140625]
[272, 304) atoms/y f64 = [0.900390625, 0.751953125, -0.751953125, -0.876953125]
[304, 336) atoms/z f64 = [0.4345703125, -0.017578125, -0.0126953125, 0.375]
The buffers are little-endian and start at multiples of 8, so a reader on a
little-endian machine can map them straight into arrays. String columns have
no buffer: their values ride in the header (element in the s values
above). A nullable column adds a mask buffer, as the ff value's torsion
parameters show (validity @…). Trajectory columns carry no precision of
their own — the collection's sequence_schema declares 0.001 for x,
y, z — but their values are already on the grid: 0.8212890625 is
841/1024.
The rules: Collection, LMDB binding (the file, frames, frame bytes).
The packed form x.mrec.zip¶
A record that is finished — no longer being appended to — can be packed into one file. Packing is a zip of the directory with every chunk copied as it is: the water record of the second section, packed:
water.mrec.zip 30 entries, all stored (method 0): true, directory entries: 0
reads back equal to water.mrec/: true
frame/atoms/_validity/b_factor/c/0 chunk 0
frame/atoms/_validity/b_factor/zarr.json array metadata
frame/atoms/_validity/zarr.json group metadata
frame/atoms/b_factor/c/0 chunk 0
frame/atoms/b_factor/zarr.json array metadata
frame/atoms/element/c/0 chunk 0
frame/atoms/element/zarr.json array metadata
frame/atoms/mass/c/0 chunk 0
frame/atoms/mass/zarr.json array metadata
frame/atoms/x/zarr.json array metadata
frame/atoms/y/c/0 chunk 0
frame/atoms/y/zarr.json array metadata
frame/atoms/z/c/0 chunk 0
frame/atoms/z/zarr.json array metadata
frame/atoms/zarr.json group metadata
frame/bonds/atomi/zarr.json array metadata
frame/bonds/atomj/c/0 chunk 0
frame/bonds/atomj/zarr.json array metadata
frame/bonds/bond_type/c/0 chunk 0
frame/bonds/bond_type/zarr.json array metadata
frame/bonds/zarr.json group metadata
frame/box/vectors/c/0/0 chunk 0/0
frame/box/vectors/zarr.json array metadata
frame/box/zarr.json group metadata
frame/density/rho/c/0 chunk 0
frame/density/rho/zarr.json array metadata
frame/density/zarr.json group metadata
frame/zarr.json group metadata
meta/zarr.json group metadata
zarr.json group metadata
- Entry names are relative to the record root:
zarr.json,meta/zarr.json,frame/atoms/x/zarr.json; the directorywater.mrec/is not part of any name. - Every entry is stored (compression method 0). Chunks are already compressed by their codec pipeline, and a second layer would cost random access.
- There are no directory entries: a Zarr store has keys, not folders.
frame/atoms/x/c/0andframe/bonds/atomi/c/0are not there: those chunks hold only the fill value (zeros), and Zarr does not write such a chunk. A reader returns the zeros.
Any zip-backed Zarr store opens it (zarr-python's ZipStore, zarrs'
ZipStorageAdapter), and reading it gives back the record exactly.
The rules: At-rest form.
Cost cheat-sheet¶
What a trajectory costs on disk follows how often its sections change,
not how long the run is. For coordinates, the bytes per atom per frame of
the three columns x, y, z through the reference pipeline (3000 atoms
scattered in a box — no molecular order, the worst case for a compressor —
over 40 frames; shard indexes not counted):
coordinates x, y, z |
pipeline | bytes per atom per frame |
|---|---|---|
lossless f64 |
bytes → crc32c | 24.0 |
precision p = 1e-3 |
bytes → shuffle(8) → zstd(3) → crc32c | 7.6 |
precision p = 1e-2 |
bytes → shuffle(8) → zstd(3) → crc32c | 5.8 |
A lossless f64 column is stored uncompressed: random low mantissa bits do
not compress, and skipping the compressor keeps reads fast. A declared
precision is what buys density: p = 1e-3 Å keeps every coordinate within
half a milliångström and stores a third of the bytes.
What the two trajectories above pay beyond their column values:
| record | section | updates | on disk beyond the column values |
|---|---|---|---|
| vibration.mrec | step |
— | attribute step_progression |
| vibration.mrec | time |
— | attribute time_progression |
| vibration.mrec | atoms |
4 | nothing: uniform_rows + dense_updates attributes |
| vibration.mrec | box |
1 | nothing: attribute vectors on box/ |
| reaction.mrec | step |
— | step array, 5 entries |
| reaction.mrec | atoms |
5 | step_index (5) + offset (6): 88 B raw |
| reaction.mrec | bonds |
2 | step_index (2) + offset (3): 40 B raw |
| reaction.mrec | atom_types |
3 | step_index (3) + offset (4): 56 B raw |
| reaction.mrec | insertions |
2 | step_index (2) + offset (3): 40 B raw |
| reaction.mrec | box |
2 | step_index (2) + vectors (2×3×3) |
Costs nothing beyond the values:
- a block whose row count is constant and which is written every frame (or once and never again): two attributes, no index;
- a fixed cell: one attribute;
- step and time numbers that are arithmetic progressions: one attribute each;
- a frame that omits a block, or presents it unchanged bit for bit: no update at all;
- topology kept in
system/: written once, whatever the run length.
Costs per change:
- every update of a block that is not regular: one
step_indexentry and oneoffsetentry (16 bytes before compression) plus the update's rows; the first irregular update also writes the index for the updates before it; - every change of the cell: one
step_indexentry and one 3×3vectorsmatrix (plusorigin/boundarywhen off their defaults); - step numbers that break the progression: the whole
steparray, 8 bytes per frame; - each per-step meta key: one value per frame.
The number of files grows with the number of arrays and with the bytes written, never with the number of frames: a shard holds thousands of chunks. In LMDB nothing is compressed — a collection is built for random reads — so coordinates cost their 24 raw bytes per atom, plus a header of a couple of hundred bytes per frame value.
The rules: Cost tracks change, Chunking and packing.