plan: amend the M0 decision record for copy-on-write; add AGENTS.md

The M0 implementation review found three of D1-D9 wrong or incomplete. The
originals stay in place with pointers to a new amendments section, so a later
session can see what changed rather than reading a rewritten history.

A1: D4 as written is unsound. The doc and overflow slabs are append-only, so
replay repairs them, but B+tree node pages are mutated in place -- after a
crash the file holds an arbitrary mix of written-back and not-written-back
pages, and once D6.3 truncates the log the data file is the only copy below
the watermark. A half-persisted tree is unrecoverable. So the checkpoint needs
shadow paging: no page below the last watermark's allocation mark is ever
stored into, and the watermark write is the atomic switch. Knock-on: node ids
cannot be page numbers, because Node.parent/next/prev are back-pointers by id
and copy-on-write would cascade; an in-RAM id->page table per index keeps every
persisted id at its current width and gives COW one pointer to fix.

A2: follows from A1 -- a page free list is a prerequisite, not the
defense-in-depth D6.2 assumed, because COW abandons every page it touches in
every epoch. The churn gate stays, retargeted at document garbage.

A3: section 5 step 4's trap was misidentified. store_record already copies
keys, so "entries must own their key bytes" is work that does not need doing.
The real problem is that a leaf record has nowhere to put a slab offset; the
resolution is to make the payload that offset, in every index, and delete
Entry.id rather than re-own it.

A4: making _id_ a unique index keys uniqueness on the canonical encode_key
rather than serialize_value, so int32 1 / int64 1 / double 1.0 collide as they
do in MongoDB -- a compatibility improvement, with a documented one-way
migration hazard for a database that already holds two such documents.

Also records that Engine.seq is never restored on open (harmless today, silent
data loss once a watermark exists), and two bugs the new spec harness found.

AGENTS.md carries the same rules into the operating guide: ground rules grow
from 7 to 9, and the old rule 6 is corrected with a note saying why.
This commit is contained in:
2026-08-03 17:08:41 +03:00
parent 86ae8fa8af
commit 13d7b79f2c
2 changed files with 702 additions and 0 deletions

198
AGENTS.md Normal file
View File

@@ -0,0 +1,198 @@
# AGENTS.md — operating guide for AI agents working in this repo
Read this first. It exists so any agent — in any session, without the
original conversation's context — can work here safely and verifiably.
## What this is
MultiforaDB: a lightweight, embedded, MongoDB-compatible document database
in Zig 0.16. It speaks the MongoDB wire protocol (OP_MSG, maxWireVersion 8)
so real drivers (`mongosh`, Node, PyMongo) connect over TCP. Everything
lives in a single-file LZ4-block log with an in-memory engine, B+tree
indexes, and TTL/unique/sparse/compound index support.
**Forward plan**: the project's direction — a full-fledged embedded,
tens-of-GB, maximally MongoDB-compatible database — is decided and written
in **[PLAN.md](PLAN.md)**. The current milestone is **M0 (mmap + WAL
storage foundation)**. Before starting any work, read PLAN.md; its decision
record (D1-D9) and ground rules are binding.
## Read first, in order
1. **PLAN.md** — the plan: decisions, milestones M0-M9, gates, work
breakdown, deferred designs. The next session always starts here.
2. **README.md** — current state, features, benchmark table, known gaps,
working-with-large-collections notes.
3. **ROADMAP.md** — design write-ups of the completed performance work
(B+tree, ordered `_id` index, compressed log, byte storage, decomposed
locks). Contains the project's ground rules and verification recipes.
## Toolchain
- **Zig 0.16.0** — pinned. Do not assume a newer Zig is compatible.
- **Node + the official `mongodb` driver** (in `tests/e2e/node_modules`)
for e2e. `npm install` in `tests/e2e` if node_modules is missing.
## Build / test / run
```sh
zig build # builds zig-out/bin/multiforadb (ReleaseFast by default)
zig build test # unit tests, ReleaseFast
zig build test -Doptimize=ReleaseSafe # unit tests with safety checks on
zig-out/bin/multiforadb --port 27017 --db data.log
```
**Critical traps:**
- **`zig build test` does NOT refresh the server binary.** After any change
under `src/`, rebuild with `zig build` before restarting the server, or
e2e will run against stale code and report failures that the source no
longer explains. (This has bitten repeatedly; it is documented in
`tests/e2e/README.md`.)
- **Debug builds are 10-200x slower** than ReleaseFast on every path. Never
benchmark (or time out waiting for) a Debug server. `zig build` defaults
to ReleaseFast on purpose.
## E2E discipline (any change to src/)
```sh
zig build # rebuild the server first!
zig build test # ReleaseFast
zig build test -Doptimize=ReleaseSafe
```
Then, against a server on port 27020 (TTL suites need `--ttl-sweep-secs 1`):
```sh
node tests/e2e/e2e.js # CRUD + operators + aggregate + errors
node tests/e2e/e2e2.js concurrent
node tests/e2e/e2e2.js crash-a # write 50 docs, kill -9 the server
node tests/e2e/e2e2.js crash-b # restart, verify all 50 survived
node tests/e2e/e2e3.js # secondary indexes
node tests/e2e/e2e4.js # TTL indexes (server must run --ttl-sweep-secs 1)
node tests/e2e/e2e6.js # self-contained full lifecycle (spawns its own server, incl. kill -9)
```
Which suites to run for a given change:
- anything touching the write path or log format → the crash pair
(e2e2 crash-a/b) and e2e6
- anything touching indexes → e2e3.js and e2e4.js
- everything → all of the above
`tests/e2e/README.md` has the full matrix, ports, and harness docs
(`big.js` for multi-GB loads, `compare-run.sh`/`bench-run.sh` for the
benchmark gate).
## Spec tests (the compatibility gate — PLAN D2)
```sh
bash tests/spec/fetch.sh # pinned mongodb/specifications suites
node tests/spec/run.js --scorecard # run + re-record tests/spec/scorecard.txt
node tests/spec/run.js --file find.json --verbose
```
`tests/spec/scorecard.txt` is committed and re-recorded every milestone (PLAN
D9), so progress is verifiable across sessions. It is a *measuring* tool: a red
baseline is the expected state at M0, `run.js` exits 0 as long as it ran, and
nothing unimplemented is ever counted as a pass. `tests/spec/README.md` has the
scope, the pins, and what is not yet asserted.
## Ground rules (binding — from ROADMAP.md and PLAN.md)
1. **Measure A/B on one harness. Do not trust the model.** Several
predictions made during the existing work were wrong in both directions.
The cheap A/B is: flip one line, rebuild, run, flip it back.
2. **Mutation-check any test guarding an invariant.** Break the thing the
test is supposed to catch and confirm it goes red before trusting it.
3. **The index invariant is absolute** (`src/index.zig` header): an index
only *generates candidates*; the full filter is re-applied afterwards.
Over-approximating is slow. Under-approximating is a wrong answer.
4. **The database must always open.** Replay never refuses to start over
recoverable damage (e.g., a unique index finding duplicates in existing
data warns and keeps going).
5. **The checkpoint never describes a state ahead of the durable log
tail** (M0 invariant): watermark advancement is the last step of the
checkpoint protocol, after the data-file fsync.
6. **No page below the stable mark is ever stored into** (M0, PLAN
Amendment A1). The whole crash story is downstream of this one rule: a
mutation that would dirty a page belonging to the last durable
checkpoint copies it to a fresh page first. Enforced structurally by a
copy-on-write page accessor, and mechanically by `mprotect`-ing the
stable prefix read-only in ReleaseSafe/test builds — a missed COW then
segfaults in the suite instead of corrupting a database silently.
7. **A checkpoint never renumbers slab offsets** (M0, Amendment A3). Index
leaves hold physical document offsets; only a full rebuild may move
documents, and it rebuilds every index in the same pass.
8. **Write-then-extend discipline for mmap**: ftruncate before touching
new pages; never fault past the end of the mapped file (SIGBUS).
9. **`Engine.seq` is restored on open** as `max(watermark.seq, max replayed
record.seq)`, and a log rewrite never emits a record whose seq is ≤ the
watermark. Today `seq` restarts at 0 on every open — harmless without a
watermark, silent data loss with one.
Note for anyone reading an older revision: rule 6 previously read "index
entries own their key bytes". That was a misdiagnosis — `store_record`
already copies keys. See PLAN Amendment A3 for what the actual trap was.
## Code style
Binding reference: **`docs/TIGER_STYLE.md`** (TigerBeetle's TigerStyle,
adopted verbatim). Read it before writing code. Its mechanical rules are
enforced: `zig fmt` clean, 4-space indent, hard 100-column line limit, no
`if` without braces unless it fits on one line, snake_case for functions/
variables/files, units last in names (`latency_ms_max`), and always-on
assertions (the repo's `assert.zig`, see below).
Project-specific rules and deliberate deviations from TigerStyle:
- Zig 0.16 idioms: `std.Io` threaded through everything, unmanaged
containers, arena-threaded allocators where the code does.
- House style: user-declared functions use `snake_case`; see existing
`src/` for the pattern. Keep `src/lib.zig` as the library root and keep
the server (`src/server.zig`) a thin front-end — the C API seam (PLAN
D1) depends on that staying clean.
- Public code invariants are asserted with the repo's own `assert.zig`
(always active, including ReleaseFast) rather than `std.debug.assert`,
so they do not silently vanish in optimized builds.
- **Deviation (allocation):** TigerStyle's "statically allocated at
startup, no allocation after init" is not yet this repo's architecture —
the engine is arena-threaded. Do not introduce new *long-lived* dynamic
allocations on hot paths; arenas scoped to a request are fine. Static
allocation is a deferred design item (PLAN).
- **Deviation (recursion):** TigerStyle bans recursion. `bson.zig`'s
document parser is still recursive (nested docs/arrays); keep new code
iterative and do not extend the recursion depth usage without a design
note. Bounding parser depth is tracked work.
- **Functions stay under 70 lines** (TigerStyle hard limit). A handful of
pre-existing offenders remain (see the reflow/refactor backlog in
ROADMAP notes); new code must not add to the count.
## Source map
## Source map
```
src/bson.zig BSON parse/serialize, ObjectId, canonical comparison/key encoding
src/wire.zig OP_MSG/OP_QUERY framing, message + reply builders
src/commands.zig command dispatch (hello, CRUD, aggregate, admin, indexes)
src/server.zig TCP accept loop, per-connection handlers, TTL sweep monitor
src/db.zig engine: db → collection → _id → document maps, slab storage
src/storage.zig append-only log: blocks, LZ4, XxHash3, replay, compaction
src/query.zig filter matcher, regex engine, sort, projection
src/index.zig B+tree indexes: entries, search, query planner
src/update.zig update operators with dot-path navigation
src/main.zig CLI: --port, --bind, --db, --ttl-sweep-secs, --compact-threshold
```
## Working on a milestone
1. Re-read PLAN.md (and this file). Confirm which milestone is current.
2. Per-milestone gates are in PLAN.md; the M0 gate is 6 items ending with
benchmark parity (phase8 baseline) and the spec-test runner red
baseline.
3. Commit scorecard and benchmark results with each milestone (PLAN D9) so
progress stays verifiable across sessions.
4. Deferred designs (cursors, aggregation, transactions, change streams,
C API) are deliberately *not* specified yet — grill the design with the
human before implementing (PLAN section 6).

504
PLAN.md Normal file
View File

@@ -0,0 +1,504 @@
# PLAN — from MVP to a full-fledged embedded MongoDB-compatible database
Status: **design agreed** (decision record below). Next action: **implement
milestone 0** (mmap + WAL storage foundation). This document is the
single source of truth for the project's direction; every milestone starts
by re-reading it. It is intentionally self-contained so work can continue
across sessions/agents without this planning conversation's context.
The current engine is documented in [README.md](README.md); the completed
performance work (B+tree, ordered `_id` index, compressed log, byte
storage, decomposed locks) is written up in [ROADMAP.md](ROADMAP.md).
Both stay current; this file adds the *forward* plan.
---
## 1. Goal and constraints
Build on the existing MVP a **full-fledged embedded document database that
is maximally MongoDB-compatible**, holding **tens of GB with excellent
performance** on ordinary machines.
- **Embedded** means: a host application runs the engine in-process; real
MongoDB drivers (`mongosh`, Node driver, PyMongo, …) connect over TCP (or
a unix socket) to a server the host starts. The engine is a library
(`src/lib.zig` already is one); `src/server.zig` is a thin front-end. A
stable C API is a later layer — the architecture must keep that seam
clean (D1).
- **Maximally compatible** is measured objectively: the official MongoDB
JSON specification test suites are the semantic gate, layered on top of
the existing hand-written e2e suites, which cover what the spec tests do
not (crash durability, compaction, TTL, lifecycle) (D2).
- **Tens of GB** forces a disk-backed engine: the current all-in-RAM model
costs ~1.5-2x data size in RSS (5 GB → 6-7 GB RAM), which is out of scope
for the target. Storage must therefore move to disk while keeping the
engine's behavior identical (D3, D4).
- **Excellent performance** is guarded by a benchmark-parity gate: no
existing benchmark row may regress against the recorded phase8 baseline
(D7).
### Out of scope for "full-fledged v1" (explicit, reconsidered only later)
- Replica sets, sharding, Atlas Search, time-series collections.
- Text and geo indexes (2d / 2dsphere). Unique / sparse / TTL / compound /
hashed / partial are the index surface.
- Full ICU collation — only the `simple` locale and a hand-rolled subset.
- OP_COMPRESSED (deferred to the last milestone; drivers negotiate and
tolerate refusal).
- Multi-process access to one data directory (like MongoDB, single-process
only; no file locking).
- GridFS needs no engine work: it is client-side and works over the wire
protocol once cursors/aggregation are in place.
---
## 2. Decision record
### D1 — Embedding model: in-process server now, architected as a library + C API later
The engine stays a library; the server is a thin front-end (177 lines
today). A host app starts the engine and runs the accept loop on its own
thread (or a unix socket). Compatibility is exercised through real drivers.
A stable C API is deliberately a *later* layer, not a current requirement —
a bespoke API can never be "MongoDB-compatible" anyway; only the wire
protocol can. The engine's public surface must stay clean enough that the
C API can be bolted on without a redesign.
### D2 — Compatibility yardstick: official spec tests + own e2e (hybrid)
- **Gate for command semantics**: the official MongoDB JSON spec tests
(github.com/mongodb/specifications) — crud, aggregate, sessions,
transactions, change streams, gridfs, collation (subset). This turns
"maximally compatible" into a concrete list of test files.
- **Runner**: Node-based (the official `mongodb` driver is already a dev
dependency). Pin both a specific commit of `mongodb/specifications` and a
specific driver version (lockfile).
- **Scorecard**: a committed file — **`tests/spec/scorecard.txt`**, with the
runner and its pin in `tests/spec/` (see `tests/spec/README.md`) — holding
per-file pass/fail/skip counts and every non-passing case with its reason,
updated every milestone. Objective progress, mirroring how benchmark
results are already recorded.
Two properties it must keep, or the number flatters the engine: nothing
unimplemented is ever counted as a pass (it is a SKIP with a reason), and
the one assertion class the runner does not yet check — `expectEvents`,
i.e. command monitoring — is disclosed at the top of the file, so `pass`
reads as an upper bound until M1 wires events up.
- **Own e2e stays**: crash pair, kill -9, compaction, TTL, concurrent
clients, lifecycle (e2e.js … e2e6.js). The spec tests do not cover these.
### D3 — Target scale: tens of GB → disk-backed engine required
All-in-RAM is ruled out at the target scale. The spec tests do not require
bounded memory — **MongoDB itself is not memory-bounded** (WiredTiger
caches ~50% of RAM) — but the *embedded* product requirement does. This
decision drives D4 and the milestone order (D5).
### D4 — Storage architecture: mmap data file + existing log as WAL
*Amended 2026-08-03 (see "Amendment A1" below): the checkpoint requires
copy-on-write. The rest of this decision stands as written.*
The data file holds the engine's structures as mmap'd regions; the
existing append-only log remains the **source of truth** (WAL) with a
**lagging checkpoint**:
- Documents live in a slab; B+tree nodes live in a flat u32-addressed
arena with an overflow slab for long keys. u32 addresses are *file
offsets*: **on-disk format == in-memory format**, no serialization on
page-in. (The ROADMAP chose the flat node array and the slab precisely so
this checkpoint story works.)
- Caching and eviction are the OS page cache's job (LMDB-style). RSS =
working set, not data size.
- The log keeps its exact record framing, torn-tail handling, group commit
and fsync-before-ack. The checkpoint may lag: on open, map the data file
and replay log records with seq > watermark. `Engine.apply_record` is
already idempotent.
- **No page below the last watermark's allocation mark is ever stored
into.** A mutation that would dirty such a page copies it to a freshly
allocated page first, and the watermark write is the atomic switch to the
new set of pages (Amendment A1). This is the premise the crash-recovery
invariant below actually rests on.
- **The in-memory docs hashmap is dropped.** At tens of GB it costs ~64-100
bytes/doc (10+ GB of RAM at 100M docs). The ordered `_id_` B+tree becomes
the primary lookup; scans become index-ordered. Trap: the B+tree leaf has
nowhere to put a document's slab offset — see Amendment A3, which
supersedes the "entries must own their key bytes" framing.
- Crash-recovery invariant, to hold absolutely:
**the checkpoint (data file) never describes a state ahead of the
durable log tail** (D6).
### D5 — Sequencing: narrow mmap foundation first, compatibility after
Milestone 0 is the riskiest technical piece (it touches every data
structure) and everything later — cursor internals (offsets vs pointers),
transactions (undo/COW), change streams (oplog tagging) — is built on top
of it. Building compatibility on the in-memory engine first would mean
writing those features twice. Therefore: **M0 = narrow mmap+WAL
foundation**, scoped tightly (below), with the spec-test runner built in
parallel (it is independent); all compatibility features land on the final
substrate.
### D6 — Foundation internals (agreed details for M0)
1. **One data file + the existing log.** Regions inside the data file:
header/metadata, B+tree node arena (+ overflow), doc slab (8 MiB
segments, as today). One file keeps backup/UX simple; growth is
ftruncate + remap (see 4).
2. **Page allocation: a free list is a prerequisite** (amended — see
Amendment A2; this item previously said "start with abandon semantics,
a freelist is defense-in-depth"). Copy-on-write abandons every node page
it touches in an epoch, *every* epoch, so a write-heavy workload grows
the file by `epochs × touched_set` without bound. The free list releases
pages with a **two-generation delay**, which also keeps generation N-1's
image intact and usable as a recovery fallback.
The **churn gate stays**, retargeted at *document* garbage in the slab:
after a 10-20 GB load with an index and ~50% churn (delete+reinsert),
the data file must be ≤ ~1.3x the live data. Dead document bytes are
reclaimed only by a full rebuild (D6.6/§5), never by a checkpoint,
because index leaves hold physical slab offsets.
3. **Watermark: double-buffered header slot** (two 4 KiB slots,
generation + checksum; pick the newer valid one) so a torn header write
can never advance the watermark past the durable log tail. Checkpoint
protocol, given the COW premise of D4 (*no page below the last
watermark's allocation mark has been stored into*): snapshot collections
(existing compaction machinery) → msync + fsync data regions → write the
catalog into freshly allocated pages, fsync → write watermark → fsync →
truncate log → advance the stable mark and rotate the free lists. The
watermark equals the group-commit sealed seq; it **never** moves past
the durable log tail. The single load-bearing ordering edge is that the
watermark write happens strictly *after* the data-page msync/fsync: the
watermark is a promise that every page it describes is on disk.
4. **File growth = ftruncate + remap.** u32 offsets are file offsets, not
pointers: even if remap moves the base, structures are unaffected.
Reserve address space for tens of GB (64-bit, MAP_NORESERVE).
5. **Single-process only** (documented; no file locking).
6. **Backup story** (a feature milestone, but the layout must not block
it): a `backup` command (checkpoint → consistent snapshot → copy),
like `sqlite3_backup`. Raw file copies are *not* consistent while the
checkpoint lags.
### D7 — Milestone 0 acceptance gate
1. Unit tests green in ReleaseFast **and** ReleaseSafe (existing
discipline).
2. All existing e2e green unchanged: e2e.js, e2e2.js (concurrent +
crash pair), e2e3.js, e2e4.js, e2e6.js — the mmap engine is a drop-in
replacement; no semantic regression.
3. Large smoke: 20-30 GB load via big.js; RSS ≈ working set, not data
size; reopen in seconds (replay from checkpoint must beat today's full
replay); kill -9 mid-write + reopen safe.
4. Churn gate (D6.2).
5. **Benchmark parity**: no phase8 row regresses (compare-run.sh +
concurrent.js) — the guardrail for "excellent performance".
6. The spec-test runner exists and runs the crud + aggregate suites
(red is fine at M0; it is the harness), scorecard file committed.
### D8 — Deferral policy
Deep design of cursors, the aggregation expression engine, transaction
snapshot isolation, and change-stream resume tokens is **deferred to their
own milestones**, when the substrate (mmap engine) is real — designs on
paper now would be for the wrong data model. Each such milestone gets its
own design review before implementation.
### D9 — Working mode
This plan is the shared context. Milestones proceed one at a time; each
starts by re-reading this document; the scorecard and benchmark results are
committed per milestone so progress is verifiable by any agent in any
session.
---
## 2.1 Amendments
Decisions D1-D9 were agreed before the M0 implementation review. That review
found three of them wrong or incomplete. The originals are left in place
above with pointers here, so a future session can see what changed and why
rather than reading a rewritten history.
### Amendment A1 — the checkpoint requires copy-on-write (amends D4, D6.3)
D4 as written is unsound. The doc slab and the overflow slab are
append-only, so replaying the log over them repairs whatever the checkpoint
missed. **B+tree node pages are mutated in place.** After a crash the file
holds an arbitrary mix of written-back and not-written-back pages; a tree
with half its dirtied pages persisted is not a tree, and replaying `upsert`
records into it does not repair it. There is no fallback either, because
D6.3's protocol truncates the log — below the watermark the data file is the
*only* copy.
So the data file must be a *consistent* snapshot at the watermark, which
forces shadow paging: **copy-on-first-touch above a stable mark, with the
watermark write as the atomic root switch.** LMDB's scheme, minus MVCC, plus
this project's WAL.
Recovery is then `image_G + replay(seq > W_G)` at every crash point, and it
holds *only* because no page in `image_G` was ever stored into. Consequence
for the hot path: **zero msyncs between checkpoints** (dirty pages may sit
in the page cache indefinitely; the kernel may write any of them back at any
time, and both outcomes are fine), so the write path keeps exactly one
fsync — the WAL's, in `Engine.commit`. Exactly one msync per checkpoint,
ordered strictly before the watermark write.
A knock-on: **node ids cannot be page numbers.** `Node.parent`, `next`,
`prev` and `Slot.extra` are back-pointers *by id*, so COWing a page would
force its siblings and its whole subtree to COW as well. A small in-RAM
`id -> page` table per index, written wholesale at each checkpoint, keeps
every persisted id at its exact current width and meaning and gives COW a
single pointer to fix. It costs one dependent load per node access and
4 B/node (~5.6 MB at 100M docs) against the 64-100 B/doc hashmap being
deleted. Dropping the back-pointers for an LMDB-style cursor path stack is
the better long-term shape and hands M1 its cursor stack for free — it is
the right post-M0 A/B, not an M0 prerequisite.
Alternative considered and rejected: a dirty flag plus a full rebuild on
unclean shutdown (no COW, no free list, much less code). Rejected because
every `kill -9` would then cost a full log replay — minutes at 20-30 GB,
against D7.3's "reopen in seconds" — and the log could never be truncated,
since a full rebuild must always remain possible.
### Amendment A2 — the page free list is a prerequisite (amends D6.2)
Follows from A1; D6.2 is rewritten in place. The "abandon semantics are
fine, a freelist is defense-in-depth" analysis was about *document*
garbage, where it is correct. COW garbage is different in kind.
### Amendment A3 — the trap is the leaf payload, not key ownership (amends D4, §5 step 4)
"When the map goes away, entries must own their key bytes" describes work
that does not need doing: `Index.store_record` already copies both the key
and the payload into the node page or the overflow slab, and `Entry.key` is
already an owned dupe. The real problem is that a leaf record has **nowhere
to put a document's slab offset** — `Slot.extra` is the payload length for a
leaf and the child node id for an internal separator.
Resolution: a leaf record becomes `key ++ offset_le`, so `extra` is always
8 and every byte-accounting site (`fits`, `record_cost`, `slot_cost`,
`balanced_cut`, `repack_keep_prefix`) is unchanged. `Entry.id` is **deleted**
rather than re-owned: every entry one document contributes shares one
payload, so the offset belongs on the mutation call, which also makes it
impossible to confuse the old offset (unique-check self-exclusion, entry
removal) with the new one. Records get *smaller*, and `lookup_eq` /
`lookup_range` come to yield `u64` values that are immune to tree mutation
instead of byte slices that alias tree pages.
Do this for **secondary** indexes too, not just `_id_`: it deletes the
`_id`-to-offset hash lookup on every index-driven query outright instead of
replacing it with a B+tree descent, and it is free on the write path because
a replace already removes and reinserts every entry in every index. The
price is the invariant now recorded in §4: a checkpoint may never renumber
slab offsets.
### Amendment A4 — `_id` uniqueness becomes canonical
`_id` uniqueness is enforced today by the docs hashmap, keyed on
`bson.serialize_value`, under which int32 `1`, int64 `1` and double `1.0`
are three distinct documents. The `_id_` B+tree is keyed on
`bson.encode_key`, which *is* canonical, so those three collide — as they do
in real MongoDB. Making `_id_` a `unique` index is therefore a
**compatibility improvement**, and it is what the guard comment at the top
of `src/index.zig` anticipated when it fenced off the old map fast path.
Migration hazard, accepted deliberately: a database written by the current
code may legitimately hold two such documents, and replay evicts by
canonical key, so one is lost on reopen. Mitigation is a loud replay warning
naming the namespace and both `_id` values, plus a documented one-way
migration. Agreed with the human rather than assumed.
---
## 3. Milestones and gates
| # | Milestone | Scope | Gate |
|---|---|---|---|
| M0 | **mmap + WAL foundation** | data file format, page/extent allocator, mmap slab + B+tree arena, copy-on-write + page free list (A1/A2), watermark/replay, checkpoint (= compaction repurposed), leaf payload → slab offset (A3), drop docs hashmap, churn measurement | D7 (6 items) |
| M1 | **Cursors + wire polish** | getMore / killCursors / batchSize; server-side cursor state with idle timeout; sessions plumbing (lsid accepted) as drivers send it; hello advertisement updates; **`moreToCome` on requests** (see the bug below); command-monitoring assertions in the spec runner | crud spec suite green; e2e green |
| M2 | **Aggregation expansion** | pipeline stages and expression engine (tiered scope defined at M2 design review) | aggregate spec suite green |
| M3 | **Update operators + index types** | $setOnInsert, $addToSet, $mul, $min/$max, $pop, $pullAll, $currentDate, pipeline updates; partial + hashed indexes | remaining crud coverage; e2e3/e2e4 green |
| M4 | **Sessions + transactions** | logical sessions, snapshot isolation on the mmap engine, write concern at commit | sessions + transactions spec suites green |
| M5 | **Change streams** | change feed + resume tokens (likely log-seq based), getMore integration | change-streams spec suite green |
| M6 | **Admin/ops commands** | dbStats, collStats, serverStatus, ping, buildInfo, listDatabases filters, dropDatabase durability (log it) | mongosh UX smoke; e2e green |
| M7 | **Auth** | SCRAM-SHA-1/256, user management commands | auth spec suite (subset) green |
| M8 | **Collation** | `simple` locale + hand-rolled subset (case-insensitive etc.) | collation spec suite (passing files) green |
| M9 | **OP_COMPRESSED** | snappy/zstd codecs, negotiation | driver handshake with compression on |
| — | **C API** (after M9 or when demanded) | stable C surface over the engine; the seam preserved since D1 | C smoke tests |
Cross-cutting in every milestone: **error-code parity** with MongoDB for
the commands touched (the spec tests assert codes), scorecard + benchmark
results committed.
### Bug found by the spec harness: `$sort` in an aggregate crashes the server
**Severity: remote, client-triggerable heap corruption. Present at `d4c9b04`.**
Fix this before anything else in M0.
In `commands.cmd_aggregate`'s `$sort` stage, the branch that materializes a
not-yet-materialized stream builds its list with the *reply arena* and then
hands it to `trees`, whose `defer trees.deinit(ctx.gpa)` frees it with the
**gpa**:
```zig
var all: std.ArrayListUnmanaged(*const bson.Document) = .empty;
for (offs.items) |off| try all.append(arena, try doc_tree(arena, coll, off));
trees.deinit(ctx.gpa);
trees = all; // arena-owned buffer, gpa-freed at scope exit
```
So the list buffer is freed by an allocator that never owned it. macOS malloc
catches it and aborts (SIGTRAP, exit 133) with no panic text, which is why the
symptom reads as "the connection closed":
```
faulting thread: mfm_free <- Allocator.rawFree
<- array_list.Aligned(*const bson.Document).deinit
<- commands.cmd_aggregate
```
Trigger: **any pipeline with `$sort` and no preceding `$group`** — e.g.
`aggregate([{$sort: {x: 1}}])`. With a `$group` first, the stream is already in
tree form, the branch is skipped, and nothing happens. Every existing e2e
aggregate case happens to sort *after* grouping, which is exactly why this
survived: `e2e.js` and `e2e6.js` both use `$match + $group + $sum + $sort`.
Fix: allocate the list with `ctx.gpa` (the `*const bson.Document` values may
stay in the arena — it outlives the command; it is only the ArrayList's own
buffer whose allocator has to match its `deinit`). Add an e2e case for a bare
`$sort` pipeline, and mutation-check it by restoring `arena` on the append.
### Bug found by the spec harness: unacknowledged writes corrupt the connection
Recorded here rather than fixed in M0, since it is M1's surface — but it is a
correctness bug, not a missing feature, and it is worth doing early because it
is a handful of lines.
`wire.Message.flags` is parsed and stored but **never read**. A driver sending
an unacknowledged write (`writeConcern: {w: 0}`) sets `moreToCome` (bit 0x2)
on the *request* and does not wait for a reply; the server replies anyway, so
that reply sits unread in the socket and every later command on that
connection reads the wrong one. Reproduced end to end:
```
acknowledged insert: ok
unacknowledged insert returned: {"acknowledged":false,"insertedId":2}
next command on same connection FAILED: MongoUnexpectedServerResponseError:
BSON element "cursor" is missing
```
`countDocuments` read the stale `insert` reply. So `w: 0` — a normal
performance choice — breaks a connection on first use, and it is invisible to
the existing e2e suites because none of them use it.
Fix: when `flags & 0x2` is set on an OP_MSG request, run the command and write
no reply. Add an e2e case for it (unacknowledged write, then a read on the same
connection), and mutation-check it by clearing the flag test.
---
## 4. Ground rules (inherited and new)
From ROADMAP.md, still binding:
- **Measure A/B on one harness. Do not trust the model.**
- **Mutation-check any test guarding an invariant** — break it and confirm
it goes red.
- **The index invariant is absolute**: an index only generates candidates;
the full filter is re-applied afterwards. Over-approximating is slow;
under-approximating is a wrong answer.
- **The database must always open**: replay never refuses to start over
recoverable damage.
New for the mmap era:
- **The checkpoint never describes a state ahead of the durable log
tail.** Watermark advancement is the last step of the checkpoint
protocol, after the data-file fsync.
- **No page below the stable mark is ever stored into.** Everything else in
the crash story is downstream of this one rule (Amendment A1). It is
enforced structurally by routing every writable page access through a
copy-on-write accessor, and mechanically by `mprotect`-ing the stable
prefix read-only in ReleaseSafe and test builds so a missed COW segfaults
in the suite instead of corrupting a database silently.
- **A checkpoint never renumbers slab offsets.** Index leaves hold physical
offsets; only a full rebuild may move documents, and it rebuilds every
index in the same pass.
- **Index entries carry a slab offset, not a document id** once the docs
map is dropped (Amendment A3).
- **Write-then-extend discipline** for mmap: ftruncate before touching new
pages; never fault past the end of the mapped file (SIGBUS protection).
- Replay idempotence of `apply_record` is a *tested invariant*, not an
assumption — it is what makes a lagging checkpoint safe.
- **`Engine.seq` is restored on open** as `max(watermark.seq, max replayed
record.seq)`. Today it unconditionally restarts at 0 (`db.zig`) and
`apply_record` ignores `record.seq` entirely; harmless without a
watermark, silent data loss with one.
- **A log rewrite never emits a record whose seq is ≤ the data file's
watermark**, or the next open discards it. Today's
`compact_snapshot_coll` stamps every re-emitted record with the live
`self.seq`; the rebuild that replaces it emits an empty log instead.
---
## 5. Milestone 0 — work breakdown (where the next session starts)
The original eight-step list is refined into the ordered commit sequence
below (same work, sequenced so that **`zig build test` in ReleaseFast *and*
ReleaseSafe plus the e2e matrix stay green at every commit**). Steps 2-5
exist to shrink step 11, which is otherwise the one unavoidably large
change.
The data file's own format is recorded in `src/pager.zig`'s header comment,
in the style of the log format at the top of `src/storage.zig`. Its shape:
an array of 4 KiB pages with a single tail-bump extent allocator — page 0
the file header, pages 1-2 the watermark double buffer, pages 3.. data. D6.1
called for a region table, but the region count is dynamic and unbounded
(one node arena per index, one slab per collection), and N contiguous
regions cannot all grow at the tail. With extents there is exactly one
growth path, therefore exactly one place where the write-then-extend
discipline lives, and `ls`/`du` stay honest for D6.6's backup story.
| # | Commit | Notes |
|---|---|---|
| 0 | amend this document | Amendments A1-A4 |
| 1 | spec-test runner + scorecard | original step 1, still fully independent |
| 2 | tighten `reserve_for`'s node bound | `n*(depth+2+n/8)+4` demands ~528 MiB of headroom for a 1000-entry multikey batch; free as ArrayList capacity, real file growth once the arena is file-backed |
| 3 | accessors for node pages and overflow bytes | pure refactor over the existing ArrayList; add the missing comptime asserts pinning `@sizeOf(Node) == 4096` and `@sizeOf(Slot)` |
| 4 | heap-allocate `Index`; move its scratch buffers out | kills the 5 KB struct memmove in `Collection.remove_index` that dangles a live `Plan.index` |
| 5 | `_id_` becomes `unique`; add an exact point lookup | Amendment A4, including the replay warning and the test-corpus updates |
| 6 | leaf payload becomes the slab offset | Amendment A3; the docs map stays as a shadow cross-check that asserts both agree |
| 7 | streaming candidates + a reverse leaf iterator | a materializing full scan is ~160 MB at the 20 GB gate; cursors are M1 but *streaming a scan* is required here |
| 8 | `src/pager.zig`: page allocator, mmap over a fixed reservation | engine-unused; unit tests only |
| 9 | watermark double buffer, catalog stream, free list | engine-unused |
| 10 | the doc slab lives in the data file | offsets become absolute; the reservation moves ahead of the log append. **First measurable point: A/B bulk-insert throughput and RSS here, before the rest depends on it** |
| 11 | node arena and overflow live in the data file | 11a overflow, 11b node pages + the id→page table |
| 12 | copy-on-write above the stable mark; two-generation free list | inert while the stable mark is 0, so testable in isolation first |
| 13 | checkpoint + watermark; open from the checkpoint; restore `Engine.seq` | log **not** yet truncated, so full replay stays a live safety net and both paths are exercised |
| 14 | truncate the log after a durable checkpoint | the safety net comes out in its own small commit |
| 15 | `compact` becomes a data-file rebuild | `compact_snapshot_coll` deleted; same tmp+rename+retry shape |
| 16 | drop the docs hashmap | original step 7 |
| 17 | gates | D7's six items; churn gate per amended D6.2 |
M0 is done when D7's six items pass. Two harness fixes are prerequisites for
the gate rather than the work: `bench-run.sh` copies its report over
`bench-latest.txt` unconditionally, including after a run that only warned,
so the baseline being defended can be clobbered; and the four B+tree dev
harnesses (`spill.zig`, `spill2.zig`, `stress.zig`, `fuzz_split.zig`) are in
no build step, so `zig build test` will not notice an API break in the one
place that fuzzes splits and >1 KB keys.
---
## 6. Deferred designs (grill each at its milestone)
- **M1 cursors**: cursor id allocation, idle expiration, batchSize
semantics, getMore against a lagging/compactable engine, cursor state
lifecycle across compaction.
- **M2 aggregation**: stage/expression tiers, which spec-test files are
the gate, whether $lookup/$unwind/facet make the first cut.
- **M4 transactions**: snapshot isolation over mmap (COW vs undo), read
concern snapshot, conflict → TransientTransactionError semantics,
retryable-writes interplay.
- **M5 change streams**: resume-token design (log seq), live fan-out,
durability of the resume point.
- **C API**: exact surface (open/close, command exec, cursor iteration,
error reporting) — after the compat milestones, on the D1 seam.