diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..79b386e --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,198 @@ +# AGENTS.md — operating guide for AI agents working in this repo + +Read this first. It exists so any agent — in any session, without the +original conversation's context — can work here safely and verifiably. + +## What this is + +MultiforaDB: a lightweight, embedded, MongoDB-compatible document database +in Zig 0.16. It speaks the MongoDB wire protocol (OP_MSG, maxWireVersion 8) +so real drivers (`mongosh`, Node, PyMongo) connect over TCP. Everything +lives in a single-file LZ4-block log with an in-memory engine, B+tree +indexes, and TTL/unique/sparse/compound index support. + +**Forward plan**: the project's direction — a full-fledged embedded, +tens-of-GB, maximally MongoDB-compatible database — is decided and written +in **[PLAN.md](PLAN.md)**. The current milestone is **M0 (mmap + WAL +storage foundation)**. Before starting any work, read PLAN.md; its decision +record (D1-D9) and ground rules are binding. + +## Read first, in order + +1. **PLAN.md** — the plan: decisions, milestones M0-M9, gates, work + breakdown, deferred designs. The next session always starts here. +2. **README.md** — current state, features, benchmark table, known gaps, + working-with-large-collections notes. +3. **ROADMAP.md** — design write-ups of the completed performance work + (B+tree, ordered `_id` index, compressed log, byte storage, decomposed + locks). Contains the project's ground rules and verification recipes. + +## Toolchain + +- **Zig 0.16.0** — pinned. Do not assume a newer Zig is compatible. +- **Node + the official `mongodb` driver** (in `tests/e2e/node_modules`) + for e2e. `npm install` in `tests/e2e` if node_modules is missing. + +## Build / test / run + +```sh +zig build # builds zig-out/bin/multiforadb (ReleaseFast by default) +zig build test # unit tests, ReleaseFast +zig build test -Doptimize=ReleaseSafe # unit tests with safety checks on +zig-out/bin/multiforadb --port 27017 --db data.log +``` + +**Critical traps:** + +- **`zig build test` does NOT refresh the server binary.** After any change + under `src/`, rebuild with `zig build` before restarting the server, or + e2e will run against stale code and report failures that the source no + longer explains. (This has bitten repeatedly; it is documented in + `tests/e2e/README.md`.) +- **Debug builds are 10-200x slower** than ReleaseFast on every path. Never + benchmark (or time out waiting for) a Debug server. `zig build` defaults + to ReleaseFast on purpose. + +## E2E discipline (any change to src/) + +```sh +zig build # rebuild the server first! +zig build test # ReleaseFast +zig build test -Doptimize=ReleaseSafe +``` + +Then, against a server on port 27020 (TTL suites need `--ttl-sweep-secs 1`): + +```sh +node tests/e2e/e2e.js # CRUD + operators + aggregate + errors +node tests/e2e/e2e2.js concurrent +node tests/e2e/e2e2.js crash-a # write 50 docs, kill -9 the server +node tests/e2e/e2e2.js crash-b # restart, verify all 50 survived +node tests/e2e/e2e3.js # secondary indexes +node tests/e2e/e2e4.js # TTL indexes (server must run --ttl-sweep-secs 1) +node tests/e2e/e2e6.js # self-contained full lifecycle (spawns its own server, incl. kill -9) +``` + +Which suites to run for a given change: + +- anything touching the write path or log format → the crash pair + (e2e2 crash-a/b) and e2e6 +- anything touching indexes → e2e3.js and e2e4.js +- everything → all of the above + +`tests/e2e/README.md` has the full matrix, ports, and harness docs +(`big.js` for multi-GB loads, `compare-run.sh`/`bench-run.sh` for the +benchmark gate). + +## Spec tests (the compatibility gate — PLAN D2) + +```sh +bash tests/spec/fetch.sh # pinned mongodb/specifications suites +node tests/spec/run.js --scorecard # run + re-record tests/spec/scorecard.txt +node tests/spec/run.js --file find.json --verbose +``` + +`tests/spec/scorecard.txt` is committed and re-recorded every milestone (PLAN +D9), so progress is verifiable across sessions. It is a *measuring* tool: a red +baseline is the expected state at M0, `run.js` exits 0 as long as it ran, and +nothing unimplemented is ever counted as a pass. `tests/spec/README.md` has the +scope, the pins, and what is not yet asserted. + +## Ground rules (binding — from ROADMAP.md and PLAN.md) + +1. **Measure A/B on one harness. Do not trust the model.** Several + predictions made during the existing work were wrong in both directions. + The cheap A/B is: flip one line, rebuild, run, flip it back. +2. **Mutation-check any test guarding an invariant.** Break the thing the + test is supposed to catch and confirm it goes red before trusting it. +3. **The index invariant is absolute** (`src/index.zig` header): an index + only *generates candidates*; the full filter is re-applied afterwards. + Over-approximating is slow. Under-approximating is a wrong answer. +4. **The database must always open.** Replay never refuses to start over + recoverable damage (e.g., a unique index finding duplicates in existing + data warns and keeps going). +5. **The checkpoint never describes a state ahead of the durable log + tail** (M0 invariant): watermark advancement is the last step of the + checkpoint protocol, after the data-file fsync. +6. **No page below the stable mark is ever stored into** (M0, PLAN + Amendment A1). The whole crash story is downstream of this one rule: a + mutation that would dirty a page belonging to the last durable + checkpoint copies it to a fresh page first. Enforced structurally by a + copy-on-write page accessor, and mechanically by `mprotect`-ing the + stable prefix read-only in ReleaseSafe/test builds — a missed COW then + segfaults in the suite instead of corrupting a database silently. +7. **A checkpoint never renumbers slab offsets** (M0, Amendment A3). Index + leaves hold physical document offsets; only a full rebuild may move + documents, and it rebuilds every index in the same pass. +8. **Write-then-extend discipline for mmap**: ftruncate before touching + new pages; never fault past the end of the mapped file (SIGBUS). +9. **`Engine.seq` is restored on open** as `max(watermark.seq, max replayed + record.seq)`, and a log rewrite never emits a record whose seq is ≤ the + watermark. Today `seq` restarts at 0 on every open — harmless without a + watermark, silent data loss with one. + +Note for anyone reading an older revision: rule 6 previously read "index +entries own their key bytes". That was a misdiagnosis — `store_record` +already copies keys. See PLAN Amendment A3 for what the actual trap was. + +## Code style + +Binding reference: **`docs/TIGER_STYLE.md`** (TigerBeetle's TigerStyle, +adopted verbatim). Read it before writing code. Its mechanical rules are +enforced: `zig fmt` clean, 4-space indent, hard 100-column line limit, no +`if` without braces unless it fits on one line, snake_case for functions/ +variables/files, units last in names (`latency_ms_max`), and always-on +assertions (the repo's `assert.zig`, see below). + +Project-specific rules and deliberate deviations from TigerStyle: + +- Zig 0.16 idioms: `std.Io` threaded through everything, unmanaged + containers, arena-threaded allocators where the code does. +- House style: user-declared functions use `snake_case`; see existing + `src/` for the pattern. Keep `src/lib.zig` as the library root and keep + the server (`src/server.zig`) a thin front-end — the C API seam (PLAN + D1) depends on that staying clean. +- Public code invariants are asserted with the repo's own `assert.zig` + (always active, including ReleaseFast) rather than `std.debug.assert`, + so they do not silently vanish in optimized builds. +- **Deviation (allocation):** TigerStyle's "statically allocated at + startup, no allocation after init" is not yet this repo's architecture — + the engine is arena-threaded. Do not introduce new *long-lived* dynamic + allocations on hot paths; arenas scoped to a request are fine. Static + allocation is a deferred design item (PLAN). +- **Deviation (recursion):** TigerStyle bans recursion. `bson.zig`'s + document parser is still recursive (nested docs/arrays); keep new code + iterative and do not extend the recursion depth usage without a design + note. Bounding parser depth is tracked work. +- **Functions stay under 70 lines** (TigerStyle hard limit). A handful of + pre-existing offenders remain (see the reflow/refactor backlog in + ROADMAP notes); new code must not add to the count. + +## Source map + +## Source map + +``` +src/bson.zig BSON parse/serialize, ObjectId, canonical comparison/key encoding +src/wire.zig OP_MSG/OP_QUERY framing, message + reply builders +src/commands.zig command dispatch (hello, CRUD, aggregate, admin, indexes) +src/server.zig TCP accept loop, per-connection handlers, TTL sweep monitor +src/db.zig engine: db → collection → _id → document maps, slab storage +src/storage.zig append-only log: blocks, LZ4, XxHash3, replay, compaction +src/query.zig filter matcher, regex engine, sort, projection +src/index.zig B+tree indexes: entries, search, query planner +src/update.zig update operators with dot-path navigation +src/main.zig CLI: --port, --bind, --db, --ttl-sweep-secs, --compact-threshold +``` + +## Working on a milestone + +1. Re-read PLAN.md (and this file). Confirm which milestone is current. +2. Per-milestone gates are in PLAN.md; the M0 gate is 6 items ending with + benchmark parity (phase8 baseline) and the spec-test runner red + baseline. +3. Commit scorecard and benchmark results with each milestone (PLAN D9) so + progress stays verifiable across sessions. +4. Deferred designs (cursors, aggregation, transactions, change streams, + C API) are deliberately *not* specified yet — grill the design with the + human before implementing (PLAN section 6). diff --git a/PLAN.md b/PLAN.md new file mode 100644 index 0000000..9f057a5 --- /dev/null +++ b/PLAN.md @@ -0,0 +1,504 @@ +# PLAN — from MVP to a full-fledged embedded MongoDB-compatible database + +Status: **design agreed** (decision record below). Next action: **implement +milestone 0** (mmap + WAL storage foundation). This document is the +single source of truth for the project's direction; every milestone starts +by re-reading it. It is intentionally self-contained so work can continue +across sessions/agents without this planning conversation's context. + +The current engine is documented in [README.md](README.md); the completed +performance work (B+tree, ordered `_id` index, compressed log, byte +storage, decomposed locks) is written up in [ROADMAP.md](ROADMAP.md). +Both stay current; this file adds the *forward* plan. + +--- + +## 1. Goal and constraints + +Build on the existing MVP a **full-fledged embedded document database that +is maximally MongoDB-compatible**, holding **tens of GB with excellent +performance** on ordinary machines. + +- **Embedded** means: a host application runs the engine in-process; real + MongoDB drivers (`mongosh`, Node driver, PyMongo, …) connect over TCP (or + a unix socket) to a server the host starts. The engine is a library + (`src/lib.zig` already is one); `src/server.zig` is a thin front-end. A + stable C API is a later layer — the architecture must keep that seam + clean (D1). +- **Maximally compatible** is measured objectively: the official MongoDB + JSON specification test suites are the semantic gate, layered on top of + the existing hand-written e2e suites, which cover what the spec tests do + not (crash durability, compaction, TTL, lifecycle) (D2). +- **Tens of GB** forces a disk-backed engine: the current all-in-RAM model + costs ~1.5-2x data size in RSS (5 GB → 6-7 GB RAM), which is out of scope + for the target. Storage must therefore move to disk while keeping the + engine's behavior identical (D3, D4). +- **Excellent performance** is guarded by a benchmark-parity gate: no + existing benchmark row may regress against the recorded phase8 baseline + (D7). + +### Out of scope for "full-fledged v1" (explicit, reconsidered only later) + +- Replica sets, sharding, Atlas Search, time-series collections. +- Text and geo indexes (2d / 2dsphere). Unique / sparse / TTL / compound / + hashed / partial are the index surface. +- Full ICU collation — only the `simple` locale and a hand-rolled subset. +- OP_COMPRESSED (deferred to the last milestone; drivers negotiate and + tolerate refusal). +- Multi-process access to one data directory (like MongoDB, single-process + only; no file locking). +- GridFS needs no engine work: it is client-side and works over the wire + protocol once cursors/aggregation are in place. + +--- + +## 2. Decision record + +### D1 — Embedding model: in-process server now, architected as a library + C API later + +The engine stays a library; the server is a thin front-end (177 lines +today). A host app starts the engine and runs the accept loop on its own +thread (or a unix socket). Compatibility is exercised through real drivers. +A stable C API is deliberately a *later* layer, not a current requirement — +a bespoke API can never be "MongoDB-compatible" anyway; only the wire +protocol can. The engine's public surface must stay clean enough that the +C API can be bolted on without a redesign. + +### D2 — Compatibility yardstick: official spec tests + own e2e (hybrid) + +- **Gate for command semantics**: the official MongoDB JSON spec tests + (github.com/mongodb/specifications) — crud, aggregate, sessions, + transactions, change streams, gridfs, collation (subset). This turns + "maximally compatible" into a concrete list of test files. +- **Runner**: Node-based (the official `mongodb` driver is already a dev + dependency). Pin both a specific commit of `mongodb/specifications` and a + specific driver version (lockfile). +- **Scorecard**: a committed file — **`tests/spec/scorecard.txt`**, with the + runner and its pin in `tests/spec/` (see `tests/spec/README.md`) — holding + per-file pass/fail/skip counts and every non-passing case with its reason, + updated every milestone. Objective progress, mirroring how benchmark + results are already recorded. + Two properties it must keep, or the number flatters the engine: nothing + unimplemented is ever counted as a pass (it is a SKIP with a reason), and + the one assertion class the runner does not yet check — `expectEvents`, + i.e. command monitoring — is disclosed at the top of the file, so `pass` + reads as an upper bound until M1 wires events up. +- **Own e2e stays**: crash pair, kill -9, compaction, TTL, concurrent + clients, lifecycle (e2e.js … e2e6.js). The spec tests do not cover these. + +### D3 — Target scale: tens of GB → disk-backed engine required + +All-in-RAM is ruled out at the target scale. The spec tests do not require +bounded memory — **MongoDB itself is not memory-bounded** (WiredTiger +caches ~50% of RAM) — but the *embedded* product requirement does. This +decision drives D4 and the milestone order (D5). + +### D4 — Storage architecture: mmap data file + existing log as WAL + +*Amended 2026-08-03 (see "Amendment A1" below): the checkpoint requires +copy-on-write. The rest of this decision stands as written.* + +The data file holds the engine's structures as mmap'd regions; the +existing append-only log remains the **source of truth** (WAL) with a +**lagging checkpoint**: + +- Documents live in a slab; B+tree nodes live in a flat u32-addressed + arena with an overflow slab for long keys. u32 addresses are *file + offsets*: **on-disk format == in-memory format**, no serialization on + page-in. (The ROADMAP chose the flat node array and the slab precisely so + this checkpoint story works.) +- Caching and eviction are the OS page cache's job (LMDB-style). RSS = + working set, not data size. +- The log keeps its exact record framing, torn-tail handling, group commit + and fsync-before-ack. The checkpoint may lag: on open, map the data file + and replay log records with seq > watermark. `Engine.apply_record` is + already idempotent. +- **No page below the last watermark's allocation mark is ever stored + into.** A mutation that would dirty such a page copies it to a freshly + allocated page first, and the watermark write is the atomic switch to the + new set of pages (Amendment A1). This is the premise the crash-recovery + invariant below actually rests on. +- **The in-memory docs hashmap is dropped.** At tens of GB it costs ~64-100 + bytes/doc (10+ GB of RAM at 100M docs). The ordered `_id_` B+tree becomes + the primary lookup; scans become index-ordered. Trap: the B+tree leaf has + nowhere to put a document's slab offset — see Amendment A3, which + supersedes the "entries must own their key bytes" framing. +- Crash-recovery invariant, to hold absolutely: + **the checkpoint (data file) never describes a state ahead of the + durable log tail** (D6). + +### D5 — Sequencing: narrow mmap foundation first, compatibility after + +Milestone 0 is the riskiest technical piece (it touches every data +structure) and everything later — cursor internals (offsets vs pointers), +transactions (undo/COW), change streams (oplog tagging) — is built on top +of it. Building compatibility on the in-memory engine first would mean +writing those features twice. Therefore: **M0 = narrow mmap+WAL +foundation**, scoped tightly (below), with the spec-test runner built in +parallel (it is independent); all compatibility features land on the final +substrate. + +### D6 — Foundation internals (agreed details for M0) + +1. **One data file + the existing log.** Regions inside the data file: + header/metadata, B+tree node arena (+ overflow), doc slab (8 MiB + segments, as today). One file keeps backup/UX simple; growth is + ftruncate + remap (see 4). +2. **Page allocation: a free list is a prerequisite** (amended — see + Amendment A2; this item previously said "start with abandon semantics, + a freelist is defense-in-depth"). Copy-on-write abandons every node page + it touches in an epoch, *every* epoch, so a write-heavy workload grows + the file by `epochs × touched_set` without bound. The free list releases + pages with a **two-generation delay**, which also keeps generation N-1's + image intact and usable as a recovery fallback. + The **churn gate stays**, retargeted at *document* garbage in the slab: + after a 10-20 GB load with an index and ~50% churn (delete+reinsert), + the data file must be ≤ ~1.3x the live data. Dead document bytes are + reclaimed only by a full rebuild (D6.6/§5), never by a checkpoint, + because index leaves hold physical slab offsets. +3. **Watermark: double-buffered header slot** (two 4 KiB slots, + generation + checksum; pick the newer valid one) so a torn header write + can never advance the watermark past the durable log tail. Checkpoint + protocol, given the COW premise of D4 (*no page below the last + watermark's allocation mark has been stored into*): snapshot collections + (existing compaction machinery) → msync + fsync data regions → write the + catalog into freshly allocated pages, fsync → write watermark → fsync → + truncate log → advance the stable mark and rotate the free lists. The + watermark equals the group-commit sealed seq; it **never** moves past + the durable log tail. The single load-bearing ordering edge is that the + watermark write happens strictly *after* the data-page msync/fsync: the + watermark is a promise that every page it describes is on disk. +4. **File growth = ftruncate + remap.** u32 offsets are file offsets, not + pointers: even if remap moves the base, structures are unaffected. + Reserve address space for tens of GB (64-bit, MAP_NORESERVE). +5. **Single-process only** (documented; no file locking). +6. **Backup story** (a feature milestone, but the layout must not block + it): a `backup` command (checkpoint → consistent snapshot → copy), + like `sqlite3_backup`. Raw file copies are *not* consistent while the + checkpoint lags. + +### D7 — Milestone 0 acceptance gate + +1. Unit tests green in ReleaseFast **and** ReleaseSafe (existing + discipline). +2. All existing e2e green unchanged: e2e.js, e2e2.js (concurrent + + crash pair), e2e3.js, e2e4.js, e2e6.js — the mmap engine is a drop-in + replacement; no semantic regression. +3. Large smoke: 20-30 GB load via big.js; RSS ≈ working set, not data + size; reopen in seconds (replay from checkpoint must beat today's full + replay); kill -9 mid-write + reopen safe. +4. Churn gate (D6.2). +5. **Benchmark parity**: no phase8 row regresses (compare-run.sh + + concurrent.js) — the guardrail for "excellent performance". +6. The spec-test runner exists and runs the crud + aggregate suites + (red is fine at M0; it is the harness), scorecard file committed. + +### D8 — Deferral policy + +Deep design of cursors, the aggregation expression engine, transaction +snapshot isolation, and change-stream resume tokens is **deferred to their +own milestones**, when the substrate (mmap engine) is real — designs on +paper now would be for the wrong data model. Each such milestone gets its +own design review before implementation. + +### D9 — Working mode + +This plan is the shared context. Milestones proceed one at a time; each +starts by re-reading this document; the scorecard and benchmark results are +committed per milestone so progress is verifiable by any agent in any +session. + +--- + +## 2.1 Amendments + +Decisions D1-D9 were agreed before the M0 implementation review. That review +found three of them wrong or incomplete. The originals are left in place +above with pointers here, so a future session can see what changed and why +rather than reading a rewritten history. + +### Amendment A1 — the checkpoint requires copy-on-write (amends D4, D6.3) + +D4 as written is unsound. The doc slab and the overflow slab are +append-only, so replaying the log over them repairs whatever the checkpoint +missed. **B+tree node pages are mutated in place.** After a crash the file +holds an arbitrary mix of written-back and not-written-back pages; a tree +with half its dirtied pages persisted is not a tree, and replaying `upsert` +records into it does not repair it. There is no fallback either, because +D6.3's protocol truncates the log — below the watermark the data file is the +*only* copy. + +So the data file must be a *consistent* snapshot at the watermark, which +forces shadow paging: **copy-on-first-touch above a stable mark, with the +watermark write as the atomic root switch.** LMDB's scheme, minus MVCC, plus +this project's WAL. + +Recovery is then `image_G + replay(seq > W_G)` at every crash point, and it +holds *only* because no page in `image_G` was ever stored into. Consequence +for the hot path: **zero msyncs between checkpoints** (dirty pages may sit +in the page cache indefinitely; the kernel may write any of them back at any +time, and both outcomes are fine), so the write path keeps exactly one +fsync — the WAL's, in `Engine.commit`. Exactly one msync per checkpoint, +ordered strictly before the watermark write. + +A knock-on: **node ids cannot be page numbers.** `Node.parent`, `next`, +`prev` and `Slot.extra` are back-pointers *by id*, so COWing a page would +force its siblings and its whole subtree to COW as well. A small in-RAM +`id -> page` table per index, written wholesale at each checkpoint, keeps +every persisted id at its exact current width and meaning and gives COW a +single pointer to fix. It costs one dependent load per node access and +4 B/node (~5.6 MB at 100M docs) against the 64-100 B/doc hashmap being +deleted. Dropping the back-pointers for an LMDB-style cursor path stack is +the better long-term shape and hands M1 its cursor stack for free — it is +the right post-M0 A/B, not an M0 prerequisite. + +Alternative considered and rejected: a dirty flag plus a full rebuild on +unclean shutdown (no COW, no free list, much less code). Rejected because +every `kill -9` would then cost a full log replay — minutes at 20-30 GB, +against D7.3's "reopen in seconds" — and the log could never be truncated, +since a full rebuild must always remain possible. + +### Amendment A2 — the page free list is a prerequisite (amends D6.2) + +Follows from A1; D6.2 is rewritten in place. The "abandon semantics are +fine, a freelist is defense-in-depth" analysis was about *document* +garbage, where it is correct. COW garbage is different in kind. + +### Amendment A3 — the trap is the leaf payload, not key ownership (amends D4, §5 step 4) + +"When the map goes away, entries must own their key bytes" describes work +that does not need doing: `Index.store_record` already copies both the key +and the payload into the node page or the overflow slab, and `Entry.key` is +already an owned dupe. The real problem is that a leaf record has **nowhere +to put a document's slab offset** — `Slot.extra` is the payload length for a +leaf and the child node id for an internal separator. + +Resolution: a leaf record becomes `key ++ offset_le`, so `extra` is always +8 and every byte-accounting site (`fits`, `record_cost`, `slot_cost`, +`balanced_cut`, `repack_keep_prefix`) is unchanged. `Entry.id` is **deleted** +rather than re-owned: every entry one document contributes shares one +payload, so the offset belongs on the mutation call, which also makes it +impossible to confuse the old offset (unique-check self-exclusion, entry +removal) with the new one. Records get *smaller*, and `lookup_eq` / +`lookup_range` come to yield `u64` values that are immune to tree mutation +instead of byte slices that alias tree pages. + +Do this for **secondary** indexes too, not just `_id_`: it deletes the +`_id`-to-offset hash lookup on every index-driven query outright instead of +replacing it with a B+tree descent, and it is free on the write path because +a replace already removes and reinserts every entry in every index. The +price is the invariant now recorded in §4: a checkpoint may never renumber +slab offsets. + +### Amendment A4 — `_id` uniqueness becomes canonical + +`_id` uniqueness is enforced today by the docs hashmap, keyed on +`bson.serialize_value`, under which int32 `1`, int64 `1` and double `1.0` +are three distinct documents. The `_id_` B+tree is keyed on +`bson.encode_key`, which *is* canonical, so those three collide — as they do +in real MongoDB. Making `_id_` a `unique` index is therefore a +**compatibility improvement**, and it is what the guard comment at the top +of `src/index.zig` anticipated when it fenced off the old map fast path. + +Migration hazard, accepted deliberately: a database written by the current +code may legitimately hold two such documents, and replay evicts by +canonical key, so one is lost on reopen. Mitigation is a loud replay warning +naming the namespace and both `_id` values, plus a documented one-way +migration. Agreed with the human rather than assumed. + +--- + +## 3. Milestones and gates + +| # | Milestone | Scope | Gate | +|---|---|---|---| +| M0 | **mmap + WAL foundation** | data file format, page/extent allocator, mmap slab + B+tree arena, copy-on-write + page free list (A1/A2), watermark/replay, checkpoint (= compaction repurposed), leaf payload → slab offset (A3), drop docs hashmap, churn measurement | D7 (6 items) | +| M1 | **Cursors + wire polish** | getMore / killCursors / batchSize; server-side cursor state with idle timeout; sessions plumbing (lsid accepted) as drivers send it; hello advertisement updates; **`moreToCome` on requests** (see the bug below); command-monitoring assertions in the spec runner | crud spec suite green; e2e green | +| M2 | **Aggregation expansion** | pipeline stages and expression engine (tiered scope defined at M2 design review) | aggregate spec suite green | +| M3 | **Update operators + index types** | $setOnInsert, $addToSet, $mul, $min/$max, $pop, $pullAll, $currentDate, pipeline updates; partial + hashed indexes | remaining crud coverage; e2e3/e2e4 green | +| M4 | **Sessions + transactions** | logical sessions, snapshot isolation on the mmap engine, write concern at commit | sessions + transactions spec suites green | +| M5 | **Change streams** | change feed + resume tokens (likely log-seq based), getMore integration | change-streams spec suite green | +| M6 | **Admin/ops commands** | dbStats, collStats, serverStatus, ping, buildInfo, listDatabases filters, dropDatabase durability (log it) | mongosh UX smoke; e2e green | +| M7 | **Auth** | SCRAM-SHA-1/256, user management commands | auth spec suite (subset) green | +| M8 | **Collation** | `simple` locale + hand-rolled subset (case-insensitive etc.) | collation spec suite (passing files) green | +| M9 | **OP_COMPRESSED** | snappy/zstd codecs, negotiation | driver handshake with compression on | +| — | **C API** (after M9 or when demanded) | stable C surface over the engine; the seam preserved since D1 | C smoke tests | + +Cross-cutting in every milestone: **error-code parity** with MongoDB for +the commands touched (the spec tests assert codes), scorecard + benchmark +results committed. + +### Bug found by the spec harness: `$sort` in an aggregate crashes the server + +**Severity: remote, client-triggerable heap corruption. Present at `d4c9b04`.** +Fix this before anything else in M0. + +In `commands.cmd_aggregate`'s `$sort` stage, the branch that materializes a +not-yet-materialized stream builds its list with the *reply arena* and then +hands it to `trees`, whose `defer trees.deinit(ctx.gpa)` frees it with the +**gpa**: + +```zig +var all: std.ArrayListUnmanaged(*const bson.Document) = .empty; +for (offs.items) |off| try all.append(arena, try doc_tree(arena, coll, off)); +trees.deinit(ctx.gpa); +trees = all; // arena-owned buffer, gpa-freed at scope exit +``` + +So the list buffer is freed by an allocator that never owned it. macOS malloc +catches it and aborts (SIGTRAP, exit 133) with no panic text, which is why the +symptom reads as "the connection closed": + +``` + faulting thread: mfm_free <- Allocator.rawFree + <- array_list.Aligned(*const bson.Document).deinit + <- commands.cmd_aggregate +``` + +Trigger: **any pipeline with `$sort` and no preceding `$group`** — e.g. +`aggregate([{$sort: {x: 1}}])`. With a `$group` first, the stream is already in +tree form, the branch is skipped, and nothing happens. Every existing e2e +aggregate case happens to sort *after* grouping, which is exactly why this +survived: `e2e.js` and `e2e6.js` both use `$match + $group + $sum + $sort`. + +Fix: allocate the list with `ctx.gpa` (the `*const bson.Document` values may +stay in the arena — it outlives the command; it is only the ArrayList's own +buffer whose allocator has to match its `deinit`). Add an e2e case for a bare +`$sort` pipeline, and mutation-check it by restoring `arena` on the append. + +### Bug found by the spec harness: unacknowledged writes corrupt the connection + +Recorded here rather than fixed in M0, since it is M1's surface — but it is a +correctness bug, not a missing feature, and it is worth doing early because it +is a handful of lines. + +`wire.Message.flags` is parsed and stored but **never read**. A driver sending +an unacknowledged write (`writeConcern: {w: 0}`) sets `moreToCome` (bit 0x2) +on the *request* and does not wait for a reply; the server replies anyway, so +that reply sits unread in the socket and every later command on that +connection reads the wrong one. Reproduced end to end: + +``` +acknowledged insert: ok +unacknowledged insert returned: {"acknowledged":false,"insertedId":2} +next command on same connection FAILED: MongoUnexpectedServerResponseError: + BSON element "cursor" is missing +``` + +`countDocuments` read the stale `insert` reply. So `w: 0` — a normal +performance choice — breaks a connection on first use, and it is invisible to +the existing e2e suites because none of them use it. + +Fix: when `flags & 0x2` is set on an OP_MSG request, run the command and write +no reply. Add an e2e case for it (unacknowledged write, then a read on the same +connection), and mutation-check it by clearing the flag test. + +--- + +## 4. Ground rules (inherited and new) + +From ROADMAP.md, still binding: + +- **Measure A/B on one harness. Do not trust the model.** +- **Mutation-check any test guarding an invariant** — break it and confirm + it goes red. +- **The index invariant is absolute**: an index only generates candidates; + the full filter is re-applied afterwards. Over-approximating is slow; + under-approximating is a wrong answer. +- **The database must always open**: replay never refuses to start over + recoverable damage. + +New for the mmap era: + +- **The checkpoint never describes a state ahead of the durable log + tail.** Watermark advancement is the last step of the checkpoint + protocol, after the data-file fsync. +- **No page below the stable mark is ever stored into.** Everything else in + the crash story is downstream of this one rule (Amendment A1). It is + enforced structurally by routing every writable page access through a + copy-on-write accessor, and mechanically by `mprotect`-ing the stable + prefix read-only in ReleaseSafe and test builds so a missed COW segfaults + in the suite instead of corrupting a database silently. +- **A checkpoint never renumbers slab offsets.** Index leaves hold physical + offsets; only a full rebuild may move documents, and it rebuilds every + index in the same pass. +- **Index entries carry a slab offset, not a document id** once the docs + map is dropped (Amendment A3). +- **Write-then-extend discipline** for mmap: ftruncate before touching new + pages; never fault past the end of the mapped file (SIGBUS protection). +- Replay idempotence of `apply_record` is a *tested invariant*, not an + assumption — it is what makes a lagging checkpoint safe. +- **`Engine.seq` is restored on open** as `max(watermark.seq, max replayed + record.seq)`. Today it unconditionally restarts at 0 (`db.zig`) and + `apply_record` ignores `record.seq` entirely; harmless without a + watermark, silent data loss with one. +- **A log rewrite never emits a record whose seq is ≤ the data file's + watermark**, or the next open discards it. Today's + `compact_snapshot_coll` stamps every re-emitted record with the live + `self.seq`; the rebuild that replaces it emits an empty log instead. + +--- + +## 5. Milestone 0 — work breakdown (where the next session starts) + +The original eight-step list is refined into the ordered commit sequence +below (same work, sequenced so that **`zig build test` in ReleaseFast *and* +ReleaseSafe plus the e2e matrix stay green at every commit**). Steps 2-5 +exist to shrink step 11, which is otherwise the one unavoidably large +change. + +The data file's own format is recorded in `src/pager.zig`'s header comment, +in the style of the log format at the top of `src/storage.zig`. Its shape: +an array of 4 KiB pages with a single tail-bump extent allocator — page 0 +the file header, pages 1-2 the watermark double buffer, pages 3.. data. D6.1 +called for a region table, but the region count is dynamic and unbounded +(one node arena per index, one slab per collection), and N contiguous +regions cannot all grow at the tail. With extents there is exactly one +growth path, therefore exactly one place where the write-then-extend +discipline lives, and `ls`/`du` stay honest for D6.6's backup story. + +| # | Commit | Notes | +|---|---|---| +| 0 | amend this document | Amendments A1-A4 | +| 1 | spec-test runner + scorecard | original step 1, still fully independent | +| 2 | tighten `reserve_for`'s node bound | `n*(depth+2+n/8)+4` demands ~528 MiB of headroom for a 1000-entry multikey batch; free as ArrayList capacity, real file growth once the arena is file-backed | +| 3 | accessors for node pages and overflow bytes | pure refactor over the existing ArrayList; add the missing comptime asserts pinning `@sizeOf(Node) == 4096` and `@sizeOf(Slot)` | +| 4 | heap-allocate `Index`; move its scratch buffers out | kills the 5 KB struct memmove in `Collection.remove_index` that dangles a live `Plan.index` | +| 5 | `_id_` becomes `unique`; add an exact point lookup | Amendment A4, including the replay warning and the test-corpus updates | +| 6 | leaf payload becomes the slab offset | Amendment A3; the docs map stays as a shadow cross-check that asserts both agree | +| 7 | streaming candidates + a reverse leaf iterator | a materializing full scan is ~160 MB at the 20 GB gate; cursors are M1 but *streaming a scan* is required here | +| 8 | `src/pager.zig`: page allocator, mmap over a fixed reservation | engine-unused; unit tests only | +| 9 | watermark double buffer, catalog stream, free list | engine-unused | +| 10 | the doc slab lives in the data file | offsets become absolute; the reservation moves ahead of the log append. **First measurable point: A/B bulk-insert throughput and RSS here, before the rest depends on it** | +| 11 | node arena and overflow live in the data file | 11a overflow, 11b node pages + the id→page table | +| 12 | copy-on-write above the stable mark; two-generation free list | inert while the stable mark is 0, so testable in isolation first | +| 13 | checkpoint + watermark; open from the checkpoint; restore `Engine.seq` | log **not** yet truncated, so full replay stays a live safety net and both paths are exercised | +| 14 | truncate the log after a durable checkpoint | the safety net comes out in its own small commit | +| 15 | `compact` becomes a data-file rebuild | `compact_snapshot_coll` deleted; same tmp+rename+retry shape | +| 16 | drop the docs hashmap | original step 7 | +| 17 | gates | D7's six items; churn gate per amended D6.2 | + +M0 is done when D7's six items pass. Two harness fixes are prerequisites for +the gate rather than the work: `bench-run.sh` copies its report over +`bench-latest.txt` unconditionally, including after a run that only warned, +so the baseline being defended can be clobbered; and the four B+tree dev +harnesses (`spill.zig`, `spill2.zig`, `stress.zig`, `fuzz_split.zig`) are in +no build step, so `zig build test` will not notice an API break in the one +place that fuzzes splits and >1 KB keys. + +--- + +## 6. Deferred designs (grill each at its milestone) + +- **M1 cursors**: cursor id allocation, idle expiration, batchSize + semantics, getMore against a lagging/compactable engine, cursor state + lifecycle across compaction. +- **M2 aggregation**: stage/expression tiers, which spec-test files are + the gate, whether $lookup/$unwind/facet make the first cut. +- **M4 transactions**: snapshot isolation over mmap (COW vs undo), read + concern snapshot, conflict → TransientTransactionError semantics, + retryable-writes interplay. +- **M5 change streams**: resume-token design (log seq), live fan-out, + durability of the resume point. +- **C API**: exact surface (open/close, command exec, cursor iteration, + error reporting) — after the compat milestones, on the D1 seam.