PLAN D2 makes the official specification suites the gate for command semantics; D7.6 asks for the harness to exist at M0 with a recorded baseline. This is that harness, pinned on both sides -- mongodb/specifications @ 615e0f9 and mongodb@7.5.0 -- because a scorecard is only comparable across milestones if a delta cannot be an upstream test change. It implements the unified format's Evaluating Matches algorithm as written, including the two rules that decide whether a pass is earned: extra keys are tolerated only in a root document, and numeric types compare flexibly. Anything unimplemented is a SKIP with a reason, never a pass, and the one assertion class not yet checked -- expectEvents, i.e. command monitoring -- is disclosed at the top of the scorecard so `pass` reads as an upper bound. First honest run: 131 pass, 161 fail, 195 skip over 175 files, zero timeouts. Getting there took four attempts, and the failures are documented in the README because each would have shipped a scorecard claiming a compatibility gap that did not exist. Two were genuine leaks in this runner (clients left open when a case timed out; clients registered for cleanup only after `await connect()`, plus abandoned cases still creating more). The third I misdiagnosed as machine load. The fourth attempt found the real cause: a leaked catalog lock in the engine, fixed separately, which alone accounts for the jump from 45 passes to 131. So the runner carries its own guards: per-operation CSOT timeouts so work is never abandoned, an active-handle census per file, an end-of-run tripwire for stray timers, a hard stop if the server dies rather than emitting hundreds of misleading ECONNREFUSED failures, and --skip/--limit for bisecting a run whose failures depend on position. The README states the rule plainly -- a long unbroken tail of timeouts is a harness bug until proven otherwise -- and the two commands that settle it. Also fixes bench-run.sh, which copied its report over bench-latest.txt unconditionally, including after a run that only warned -- so a degraded run could silently replace the baseline that PLAN D7.5 makes a milestone gate.
559 lines
31 KiB
Markdown
559 lines
31 KiB
Markdown
# PLAN — from MVP to a full-fledged embedded MongoDB-compatible database
|
||
|
||
Status: **design agreed** (decision record below). Next action: **implement
|
||
milestone 0** (mmap + WAL storage foundation). This document is the
|
||
single source of truth for the project's direction; every milestone starts
|
||
by re-reading it. It is intentionally self-contained so work can continue
|
||
across sessions/agents without this planning conversation's context.
|
||
|
||
The current engine is documented in [README.md](README.md); the completed
|
||
performance work (B+tree, ordered `_id` index, compressed log, byte
|
||
storage, decomposed locks) is written up in [ROADMAP.md](ROADMAP.md).
|
||
Both stay current; this file adds the *forward* plan.
|
||
|
||
---
|
||
|
||
## 1. Goal and constraints
|
||
|
||
Build on the existing MVP a **full-fledged embedded document database that
|
||
is maximally MongoDB-compatible**, holding **tens of GB with excellent
|
||
performance** on ordinary machines.
|
||
|
||
- **Embedded** means: a host application runs the engine in-process; real
|
||
MongoDB drivers (`mongosh`, Node driver, PyMongo, …) connect over TCP (or
|
||
a unix socket) to a server the host starts. The engine is a library
|
||
(`src/lib.zig` already is one); `src/server.zig` is a thin front-end. A
|
||
stable C API is a later layer — the architecture must keep that seam
|
||
clean (D1).
|
||
- **Maximally compatible** is measured objectively: the official MongoDB
|
||
JSON specification test suites are the semantic gate, layered on top of
|
||
the existing hand-written e2e suites, which cover what the spec tests do
|
||
not (crash durability, compaction, TTL, lifecycle) (D2).
|
||
- **Tens of GB** forces a disk-backed engine: the current all-in-RAM model
|
||
costs ~1.5-2x data size in RSS (5 GB → 6-7 GB RAM), which is out of scope
|
||
for the target. Storage must therefore move to disk while keeping the
|
||
engine's behavior identical (D3, D4).
|
||
- **Excellent performance** is guarded by a benchmark-parity gate: no
|
||
existing benchmark row may regress against the recorded phase8 baseline
|
||
(D7).
|
||
|
||
### Out of scope for "full-fledged v1" (explicit, reconsidered only later)
|
||
|
||
- Replica sets, sharding, Atlas Search, time-series collections.
|
||
- Text and geo indexes (2d / 2dsphere). Unique / sparse / TTL / compound /
|
||
hashed / partial are the index surface.
|
||
- Full ICU collation — only the `simple` locale and a hand-rolled subset.
|
||
- OP_COMPRESSED (deferred to the last milestone; drivers negotiate and
|
||
tolerate refusal).
|
||
- Multi-process access to one data directory (like MongoDB, single-process
|
||
only; no file locking).
|
||
- GridFS needs no engine work: it is client-side and works over the wire
|
||
protocol once cursors/aggregation are in place.
|
||
|
||
---
|
||
|
||
## 2. Decision record
|
||
|
||
### D1 — Embedding model: in-process server now, architected as a library + C API later
|
||
|
||
The engine stays a library; the server is a thin front-end (177 lines
|
||
today). A host app starts the engine and runs the accept loop on its own
|
||
thread (or a unix socket). Compatibility is exercised through real drivers.
|
||
A stable C API is deliberately a *later* layer, not a current requirement —
|
||
a bespoke API can never be "MongoDB-compatible" anyway; only the wire
|
||
protocol can. The engine's public surface must stay clean enough that the
|
||
C API can be bolted on without a redesign.
|
||
|
||
### D2 — Compatibility yardstick: official spec tests + own e2e (hybrid)
|
||
|
||
- **Gate for command semantics**: the official MongoDB JSON spec tests
|
||
(github.com/mongodb/specifications) — crud, aggregate, sessions,
|
||
transactions, change streams, gridfs, collation (subset). This turns
|
||
"maximally compatible" into a concrete list of test files.
|
||
- **Runner**: Node-based (the official `mongodb` driver is already a dev
|
||
dependency). Pin both a specific commit of `mongodb/specifications` and a
|
||
specific driver version (lockfile).
|
||
- **Scorecard**: a committed file — **`tests/spec/scorecard.txt`**, with the
|
||
runner and its pin in `tests/spec/` (see `tests/spec/README.md`) — holding
|
||
per-file pass/fail/skip counts and every non-passing case with its reason,
|
||
updated every milestone. Objective progress, mirroring how benchmark
|
||
results are already recorded.
|
||
Two properties it must keep, or the number flatters the engine: nothing
|
||
unimplemented is ever counted as a pass (it is a SKIP with a reason), and
|
||
the one assertion class the runner does not yet check — `expectEvents`,
|
||
i.e. command monitoring — is disclosed at the top of the file, so `pass`
|
||
reads as an upper bound until M1 wires events up.
|
||
- **Own e2e stays**: crash pair, kill -9, compaction, TTL, concurrent
|
||
clients, lifecycle (e2e.js … e2e6.js). The spec tests do not cover these.
|
||
|
||
### D3 — Target scale: tens of GB → disk-backed engine required
|
||
|
||
All-in-RAM is ruled out at the target scale. The spec tests do not require
|
||
bounded memory — **MongoDB itself is not memory-bounded** (WiredTiger
|
||
caches ~50% of RAM) — but the *embedded* product requirement does. This
|
||
decision drives D4 and the milestone order (D5).
|
||
|
||
### D4 — Storage architecture: mmap data file + existing log as WAL
|
||
|
||
*Amended 2026-08-03 (see "Amendment A1" below): the checkpoint requires
|
||
copy-on-write. The rest of this decision stands as written.*
|
||
|
||
The data file holds the engine's structures as mmap'd regions; the
|
||
existing append-only log remains the **source of truth** (WAL) with a
|
||
**lagging checkpoint**:
|
||
|
||
- Documents live in a slab; B+tree nodes live in a flat u32-addressed
|
||
arena with an overflow slab for long keys. u32 addresses are *file
|
||
offsets*: **on-disk format == in-memory format**, no serialization on
|
||
page-in. (The ROADMAP chose the flat node array and the slab precisely so
|
||
this checkpoint story works.)
|
||
- Caching and eviction are the OS page cache's job (LMDB-style). RSS =
|
||
working set, not data size.
|
||
- The log keeps its exact record framing, torn-tail handling, group commit
|
||
and fsync-before-ack. The checkpoint may lag: on open, map the data file
|
||
and replay log records with seq > watermark. `Engine.apply_record` is
|
||
already idempotent.
|
||
- **No page below the last watermark's allocation mark is ever stored
|
||
into.** A mutation that would dirty such a page copies it to a freshly
|
||
allocated page first, and the watermark write is the atomic switch to the
|
||
new set of pages (Amendment A1). This is the premise the crash-recovery
|
||
invariant below actually rests on.
|
||
- **The in-memory docs hashmap is dropped.** At tens of GB it costs ~64-100
|
||
bytes/doc (10+ GB of RAM at 100M docs). The ordered `_id_` B+tree becomes
|
||
the primary lookup; scans become index-ordered. Trap: the B+tree leaf has
|
||
nowhere to put a document's slab offset — see Amendment A3, which
|
||
supersedes the "entries must own their key bytes" framing.
|
||
- Crash-recovery invariant, to hold absolutely:
|
||
**the checkpoint (data file) never describes a state ahead of the
|
||
durable log tail** (D6).
|
||
|
||
### D5 — Sequencing: narrow mmap foundation first, compatibility after
|
||
|
||
Milestone 0 is the riskiest technical piece (it touches every data
|
||
structure) and everything later — cursor internals (offsets vs pointers),
|
||
transactions (undo/COW), change streams (oplog tagging) — is built on top
|
||
of it. Building compatibility on the in-memory engine first would mean
|
||
writing those features twice. Therefore: **M0 = narrow mmap+WAL
|
||
foundation**, scoped tightly (below), with the spec-test runner built in
|
||
parallel (it is independent); all compatibility features land on the final
|
||
substrate.
|
||
|
||
### D6 — Foundation internals (agreed details for M0)
|
||
|
||
1. **One data file + the existing log.** Regions inside the data file:
|
||
header/metadata, B+tree node arena (+ overflow), doc slab (8 MiB
|
||
segments, as today). One file keeps backup/UX simple; growth is
|
||
ftruncate + remap (see 4).
|
||
2. **Page allocation: a free list is a prerequisite** (amended — see
|
||
Amendment A2; this item previously said "start with abandon semantics,
|
||
a freelist is defense-in-depth"). Copy-on-write abandons every node page
|
||
it touches in an epoch, *every* epoch, so a write-heavy workload grows
|
||
the file by `epochs × touched_set` without bound. The free list releases
|
||
pages with a **two-generation delay**, which also keeps generation N-1's
|
||
image intact and usable as a recovery fallback.
|
||
The **churn gate stays**, retargeted at *document* garbage in the slab:
|
||
after a 10-20 GB load with an index and ~50% churn (delete+reinsert),
|
||
the data file must be ≤ ~1.3x the live data. Dead document bytes are
|
||
reclaimed only by a full rebuild (D6.6/§5), never by a checkpoint,
|
||
because index leaves hold physical slab offsets.
|
||
3. **Watermark: double-buffered header slot** (two 4 KiB slots,
|
||
generation + checksum; pick the newer valid one) so a torn header write
|
||
can never advance the watermark past the durable log tail. Checkpoint
|
||
protocol, given the COW premise of D4 (*no page below the last
|
||
watermark's allocation mark has been stored into*): snapshot collections
|
||
(existing compaction machinery) → msync + fsync data regions → write the
|
||
catalog into freshly allocated pages, fsync → write watermark → fsync →
|
||
truncate log → advance the stable mark and rotate the free lists. The
|
||
watermark equals the group-commit sealed seq; it **never** moves past
|
||
the durable log tail. The single load-bearing ordering edge is that the
|
||
watermark write happens strictly *after* the data-page msync/fsync: the
|
||
watermark is a promise that every page it describes is on disk.
|
||
4. **File growth = ftruncate + remap.** u32 offsets are file offsets, not
|
||
pointers: even if remap moves the base, structures are unaffected.
|
||
Reserve address space for tens of GB (64-bit, MAP_NORESERVE).
|
||
5. **Single-process only** (documented; no file locking).
|
||
6. **Backup story** (a feature milestone, but the layout must not block
|
||
it): a `backup` command (checkpoint → consistent snapshot → copy),
|
||
like `sqlite3_backup`. Raw file copies are *not* consistent while the
|
||
checkpoint lags.
|
||
|
||
### D7 — Milestone 0 acceptance gate
|
||
|
||
1. Unit tests green in ReleaseFast **and** ReleaseSafe (existing
|
||
discipline).
|
||
2. All existing e2e green unchanged: e2e.js, e2e2.js (concurrent +
|
||
crash pair), e2e3.js, e2e4.js, e2e6.js — the mmap engine is a drop-in
|
||
replacement; no semantic regression.
|
||
3. Large smoke: 20-30 GB load via big.js; RSS ≈ working set, not data
|
||
size; reopen in seconds (replay from checkpoint must beat today's full
|
||
replay); kill -9 mid-write + reopen safe.
|
||
4. Churn gate (D6.2).
|
||
5. **Benchmark parity**: no phase8 row regresses (compare-run.sh +
|
||
concurrent.js) — the guardrail for "excellent performance".
|
||
6. The spec-test runner exists and runs the crud + aggregate suites
|
||
(red is fine at M0; it is the harness), scorecard file committed.
|
||
|
||
### D8 — Deferral policy
|
||
|
||
Deep design of cursors, the aggregation expression engine, transaction
|
||
snapshot isolation, and change-stream resume tokens is **deferred to their
|
||
own milestones**, when the substrate (mmap engine) is real — designs on
|
||
paper now would be for the wrong data model. Each such milestone gets its
|
||
own design review before implementation.
|
||
|
||
### D9 — Working mode
|
||
|
||
This plan is the shared context. Milestones proceed one at a time; each
|
||
starts by re-reading this document; the scorecard and benchmark results are
|
||
committed per milestone so progress is verifiable by any agent in any
|
||
session.
|
||
|
||
---
|
||
|
||
## 2.1 Amendments
|
||
|
||
Decisions D1-D9 were agreed before the M0 implementation review. That review
|
||
found three of them wrong or incomplete. The originals are left in place
|
||
above with pointers here, so a future session can see what changed and why
|
||
rather than reading a rewritten history.
|
||
|
||
### Amendment A1 — the checkpoint requires copy-on-write (amends D4, D6.3)
|
||
|
||
D4 as written is unsound. The doc slab and the overflow slab are
|
||
append-only, so replaying the log over them repairs whatever the checkpoint
|
||
missed. **B+tree node pages are mutated in place.** After a crash the file
|
||
holds an arbitrary mix of written-back and not-written-back pages; a tree
|
||
with half its dirtied pages persisted is not a tree, and replaying `upsert`
|
||
records into it does not repair it. There is no fallback either, because
|
||
D6.3's protocol truncates the log — below the watermark the data file is the
|
||
*only* copy.
|
||
|
||
So the data file must be a *consistent* snapshot at the watermark, which
|
||
forces shadow paging: **copy-on-first-touch above a stable mark, with the
|
||
watermark write as the atomic root switch.** LMDB's scheme, minus MVCC, plus
|
||
this project's WAL.
|
||
|
||
Recovery is then `image_G + replay(seq > W_G)` at every crash point, and it
|
||
holds *only* because no page in `image_G` was ever stored into. Consequence
|
||
for the hot path: **zero msyncs between checkpoints** (dirty pages may sit
|
||
in the page cache indefinitely; the kernel may write any of them back at any
|
||
time, and both outcomes are fine), so the write path keeps exactly one
|
||
fsync — the WAL's, in `Engine.commit`. Exactly one msync per checkpoint,
|
||
ordered strictly before the watermark write.
|
||
|
||
A knock-on: **node ids cannot be page numbers.** `Node.parent`, `next`,
|
||
`prev` and `Slot.extra` are back-pointers *by id*, so COWing a page would
|
||
force its siblings and its whole subtree to COW as well. A small in-RAM
|
||
`id -> page` table per index, written wholesale at each checkpoint, keeps
|
||
every persisted id at its exact current width and meaning and gives COW a
|
||
single pointer to fix. It costs one dependent load per node access and
|
||
4 B/node (~5.6 MB at 100M docs) against the 64-100 B/doc hashmap being
|
||
deleted. Dropping the back-pointers for an LMDB-style cursor path stack is
|
||
the better long-term shape and hands M1 its cursor stack for free — it is
|
||
the right post-M0 A/B, not an M0 prerequisite.
|
||
|
||
Alternative considered and rejected: a dirty flag plus a full rebuild on
|
||
unclean shutdown (no COW, no free list, much less code). Rejected because
|
||
every `kill -9` would then cost a full log replay — minutes at 20-30 GB,
|
||
against D7.3's "reopen in seconds" — and the log could never be truncated,
|
||
since a full rebuild must always remain possible.
|
||
|
||
### Amendment A2 — the page free list is a prerequisite (amends D6.2)
|
||
|
||
Follows from A1; D6.2 is rewritten in place. The "abandon semantics are
|
||
fine, a freelist is defense-in-depth" analysis was about *document*
|
||
garbage, where it is correct. COW garbage is different in kind.
|
||
|
||
### Amendment A3 — the trap is the leaf payload, not key ownership (amends D4, §5 step 4)
|
||
|
||
"When the map goes away, entries must own their key bytes" describes work
|
||
that does not need doing: `Index.store_record` already copies both the key
|
||
and the payload into the node page or the overflow slab, and `Entry.key` is
|
||
already an owned dupe. The real problem is that a leaf record has **nowhere
|
||
to put a document's slab offset** — `Slot.extra` is the payload length for a
|
||
leaf and the child node id for an internal separator.
|
||
|
||
Resolution: a leaf record becomes `key ++ offset_le`, so `extra` is always
|
||
8 and every byte-accounting site (`fits`, `record_cost`, `slot_cost`,
|
||
`balanced_cut`, `repack_keep_prefix`) is unchanged. `Entry.id` is **deleted**
|
||
rather than re-owned: every entry one document contributes shares one
|
||
payload, so the offset belongs on the mutation call, which also makes it
|
||
impossible to confuse the old offset (unique-check self-exclusion, entry
|
||
removal) with the new one. Records get *smaller*, and `lookup_eq` /
|
||
`lookup_range` come to yield `u64` values that are immune to tree mutation
|
||
instead of byte slices that alias tree pages.
|
||
|
||
Do this for **secondary** indexes too, not just `_id_`: it deletes the
|
||
`_id`-to-offset hash lookup on every index-driven query outright instead of
|
||
replacing it with a B+tree descent, and it is free on the write path because
|
||
a replace already removes and reinserts every entry in every index. The
|
||
price is the invariant now recorded in §4: a checkpoint may never renumber
|
||
slab offsets.
|
||
|
||
### Amendment A4 — `_id` uniqueness becomes canonical
|
||
|
||
`_id` uniqueness is enforced today by the docs hashmap, keyed on
|
||
`bson.serialize_value`, under which int32 `1`, int64 `1` and double `1.0`
|
||
are three distinct documents. The `_id_` B+tree is keyed on
|
||
`bson.encode_key`, which *is* canonical, so those three collide — as they do
|
||
in real MongoDB. Making `_id_` a `unique` index is therefore a
|
||
**compatibility improvement**, and it is what the guard comment at the top
|
||
of `src/index.zig` anticipated when it fenced off the old map fast path.
|
||
|
||
Migration hazard, accepted deliberately: a database written by the current
|
||
code may legitimately hold two such documents. Agreed with the human rather
|
||
than assumed.
|
||
|
||
*Corrected while implementing:* the hazard does **not** bite when `_id_`
|
||
becomes unique, and the reason matters for sequencing. Replay does not
|
||
maintain `_id_` entries — `apply_record` evicts through the docs map, keyed on
|
||
`serialize_value`, and the tree is bulk-built afterwards by
|
||
`build_all_indexes`. That build calls `finish_bulk` with enforcement off,
|
||
which tolerates duplicate keys and already warns ("unique index '_id_' has
|
||
duplicate keys in existing data"), honouring the "database must always open"
|
||
rule. So such a database reopens with both documents intact and a loud
|
||
warning, and new colliding inserts are rejected from then on.
|
||
|
||
The document *loss* arrives only with the commit that drops the docs hashmap,
|
||
because eviction then goes through the canonical tree and the second document
|
||
takes the first one's place. That commit is where this needs handling — a
|
||
pre-flight scan for compare-equal `_id`s, refusing to drop the map silently
|
||
while any exist — not here.
|
||
|
||
---
|
||
|
||
## 3. Milestones and gates
|
||
|
||
| # | Milestone | Scope | Gate |
|
||
|---|---|---|---|
|
||
| M0 | **mmap + WAL foundation** | data file format, page/extent allocator, mmap slab + B+tree arena, copy-on-write + page free list (A1/A2), watermark/replay, checkpoint (= compaction repurposed), leaf payload → slab offset (A3), drop docs hashmap, churn measurement | D7 (6 items) |
|
||
| M1 | **Cursors + wire polish** | getMore / killCursors / batchSize; server-side cursor state with idle timeout; sessions plumbing (lsid accepted) as drivers send it; hello advertisement updates; **`moreToCome` on requests** (see the bug below); command-monitoring assertions in the spec runner | crud spec suite green; e2e green |
|
||
| M2 | **Aggregation expansion** | pipeline stages and expression engine (tiered scope defined at M2 design review) | aggregate spec suite green |
|
||
| M3 | **Update operators + index types** | $setOnInsert, $addToSet, $mul, $min/$max, $pop, $pullAll, $currentDate, pipeline updates; partial + hashed indexes | remaining crud coverage; e2e3/e2e4 green |
|
||
| M4 | **Sessions + transactions** | logical sessions, snapshot isolation on the mmap engine, write concern at commit | sessions + transactions spec suites green |
|
||
| M5 | **Change streams** | change feed + resume tokens (likely log-seq based), getMore integration | change-streams spec suite green |
|
||
| M6 | **Admin/ops commands** | dbStats, collStats, serverStatus, ping, buildInfo, listDatabases filters, dropDatabase durability (log it) | mongosh UX smoke; e2e green |
|
||
| M7 | **Auth** | SCRAM-SHA-1/256, user management commands | auth spec suite (subset) green |
|
||
| M8 | **Collation** | `simple` locale + hand-rolled subset (case-insensitive etc.) | collation spec suite (passing files) green |
|
||
| M9 | **OP_COMPRESSED** | snappy/zstd codecs, negotiation | driver handshake with compression on |
|
||
| — | **C API** (after M9 or when demanded) | stable C surface over the engine; the seam preserved since D1 | C smoke tests |
|
||
|
||
Cross-cutting in every milestone: **error-code parity** with MongoDB for
|
||
the commands touched (the spec tests assert codes), scorecard + benchmark
|
||
results committed.
|
||
|
||
### Bug found by the spec harness: `$sort` in an aggregate crashes the server
|
||
|
||
**Severity: remote, client-triggerable heap corruption. Present at `d4c9b04`.**
|
||
Fix this before anything else in M0.
|
||
|
||
In `commands.cmd_aggregate`'s `$sort` stage, the branch that materializes a
|
||
not-yet-materialized stream builds its list with the *reply arena* and then
|
||
hands it to `trees`, whose `defer trees.deinit(ctx.gpa)` frees it with the
|
||
**gpa**:
|
||
|
||
```zig
|
||
var all: std.ArrayListUnmanaged(*const bson.Document) = .empty;
|
||
for (offs.items) |off| try all.append(arena, try doc_tree(arena, coll, off));
|
||
trees.deinit(ctx.gpa);
|
||
trees = all; // arena-owned buffer, gpa-freed at scope exit
|
||
```
|
||
|
||
So the list buffer is freed by an allocator that never owned it. macOS malloc
|
||
catches it and aborts (SIGTRAP, exit 133) with no panic text, which is why the
|
||
symptom reads as "the connection closed":
|
||
|
||
```
|
||
faulting thread: mfm_free <- Allocator.rawFree
|
||
<- array_list.Aligned(*const bson.Document).deinit
|
||
<- commands.cmd_aggregate
|
||
```
|
||
|
||
Trigger: **any pipeline with `$sort` and no preceding `$group`** — e.g.
|
||
`aggregate([{$sort: {x: 1}}])`. With a `$group` first, the stream is already in
|
||
tree form, the branch is skipped, and nothing happens. Every existing e2e
|
||
aggregate case happens to sort *after* grouping, which is exactly why this
|
||
survived: `e2e.js` and `e2e6.js` both use `$match + $group + $sum + $sort`.
|
||
|
||
Fix: allocate the list with `ctx.gpa` (the `*const bson.Document` values may
|
||
stay in the arena — it outlives the command; it is only the ArrayList's own
|
||
buffer whose allocator has to match its `deinit`). Add an e2e case for a bare
|
||
`$sort` pipeline, and mutation-check it by restoring `arena` on the append.
|
||
|
||
### Bug found by the spec harness: a nameless command leaked the catalog lock
|
||
|
||
**Fixed. Severity: permanent denial of service, remotely triggerable by an
|
||
ordinary query.** This one blocked the M0 scorecard outright and cost three
|
||
invalid baselines.
|
||
|
||
`dispatch` resolved the namespace *after* taking the catalog lock and bailed
|
||
with `orelse return` when the collection name was missing. A plain return runs
|
||
neither the `errdefer` nor the explicit unlocks, so the lock was held — shared —
|
||
for the life of the process. `db.aggregate(...)` reaches it: that sends
|
||
`{aggregate: 1}`, whose value is not a string.
|
||
|
||
Worth understanding how it hid, because the shape recurs: a leaked *shared* lock
|
||
is invisible to readers. `ping` and `listDatabases` kept answering in
|
||
microseconds and an external prober saw `ok 15ms` straight through the hang, so
|
||
the server looked healthy. Only a write needing the catalog exclusive to create
|
||
a collection blocked — so the damage appeared one command later, on a different
|
||
connection, as a client-side timeout with nothing linking it to the cause.
|
||
|
||
What actually found it: the driver's own command log, showing an insert sitting
|
||
for exactly `socketTimeoutMS` against an idle engine.
|
||
|
||
Namespace resolution now happens before any lock is taken. Two lessons kept in
|
||
`tests/spec/README.md`: a responsive server does not exonerate the engine, and
|
||
probe with the operation that is stuck rather than with `ping`.
|
||
|
||
### Bug found by the spec harness: unacknowledged writes corrupt the connection
|
||
|
||
**Fixed** (in M0 rather than M1 as originally recorded — it is a correctness
|
||
bug, not a missing feature, and it is a handful of lines).
|
||
|
||
`wire.Message.flags` is parsed and stored but **never read**. A driver sending
|
||
an unacknowledged write (`writeConcern: {w: 0}`) sets `moreToCome` (bit 0x2)
|
||
on the *request* and does not wait for a reply; the server replies anyway, so
|
||
that reply sits unread in the socket and every later command on that
|
||
connection reads the wrong one. Reproduced end to end:
|
||
|
||
```
|
||
acknowledged insert: ok
|
||
unacknowledged insert returned: {"acknowledged":false,"insertedId":2}
|
||
next command on same connection FAILED: MongoUnexpectedServerResponseError:
|
||
BSON element "cursor" is missing
|
||
```
|
||
|
||
`countDocuments` read the stale `insert` reply. So `w: 0` — a normal
|
||
performance choice — breaks a connection on first use, and it is invisible to
|
||
the existing e2e suites because none of them use it.
|
||
|
||
Fixed by suppressing the reply when `flags & 0x2` is set on an OP_MSG request:
|
||
the command still runs. The e2e case pins `maxPoolSize` to 1, since with a
|
||
larger pool the driver may hand the next operation a different connection and
|
||
hide it.
|
||
|
||
### What the harness was worth
|
||
|
||
Its first honest run — after the two bugs above and two genuine leaks in the
|
||
runner itself — reports **131 pass, 161 fail, 195 skip** over 175 files, with
|
||
zero timeouts. Before the catalog-lock fix the same suite reported 45 passes:
|
||
that gap is the measure of what one leaked lock was hiding, and of why a
|
||
scorecard is only worth committing once it disagrees with nothing that passes in
|
||
isolation.
|
||
|
||
The remaining failures are real work, and they cluster usefully: update-operator
|
||
gaps (`bad update`, `update must be a document` — M3), unimplemented commands
|
||
(`distinct`, `$merge`, `$out` — M2/M3), and result-shape mismatches in
|
||
`bulkWrite`/`insertMany`. That list, not the total, is the milestone backlog.
|
||
|
||
---
|
||
|
||
## 4. Ground rules (inherited and new)
|
||
|
||
From ROADMAP.md, still binding:
|
||
|
||
- **Measure A/B on one harness. Do not trust the model.**
|
||
- **Mutation-check any test guarding an invariant** — break it and confirm
|
||
it goes red.
|
||
- **The index invariant is absolute**: an index only generates candidates;
|
||
the full filter is re-applied afterwards. Over-approximating is slow;
|
||
under-approximating is a wrong answer.
|
||
- **The database must always open**: replay never refuses to start over
|
||
recoverable damage.
|
||
|
||
New for the mmap era:
|
||
|
||
- **The checkpoint never describes a state ahead of the durable log
|
||
tail.** Watermark advancement is the last step of the checkpoint
|
||
protocol, after the data-file fsync.
|
||
- **No page below the stable mark is ever stored into.** Everything else in
|
||
the crash story is downstream of this one rule (Amendment A1). It is
|
||
enforced structurally by routing every writable page access through a
|
||
copy-on-write accessor, and mechanically by `mprotect`-ing the stable
|
||
prefix read-only in ReleaseSafe and test builds so a missed COW segfaults
|
||
in the suite instead of corrupting a database silently.
|
||
- **A checkpoint never renumbers slab offsets.** Index leaves hold physical
|
||
offsets; only a full rebuild may move documents, and it rebuilds every
|
||
index in the same pass.
|
||
- **Index entries carry a slab offset, not a document id** once the docs
|
||
map is dropped (Amendment A3).
|
||
- **Write-then-extend discipline** for mmap: ftruncate before touching new
|
||
pages; never fault past the end of the mapped file (SIGBUS protection).
|
||
- Replay idempotence of `apply_record` is a *tested invariant*, not an
|
||
assumption — it is what makes a lagging checkpoint safe.
|
||
- **`Engine.seq` is restored on open** as `max(watermark.seq, max replayed
|
||
record.seq)`. Today it unconditionally restarts at 0 (`db.zig`) and
|
||
`apply_record` ignores `record.seq` entirely; harmless without a
|
||
watermark, silent data loss with one.
|
||
- **A log rewrite never emits a record whose seq is ≤ the data file's
|
||
watermark**, or the next open discards it. Today's
|
||
`compact_snapshot_coll` stamps every re-emitted record with the live
|
||
`self.seq`; the rebuild that replaces it emits an empty log instead.
|
||
|
||
---
|
||
|
||
## 5. Milestone 0 — work breakdown (where the next session starts)
|
||
|
||
The original eight-step list is refined into the ordered commit sequence
|
||
below (same work, sequenced so that **`zig build test` in ReleaseFast *and*
|
||
ReleaseSafe plus the e2e matrix stay green at every commit**). Steps 2-5
|
||
exist to shrink step 11, which is otherwise the one unavoidably large
|
||
change.
|
||
|
||
The data file's own format is recorded in `src/pager.zig`'s header comment,
|
||
in the style of the log format at the top of `src/storage.zig`. Its shape:
|
||
an array of 4 KiB pages with a single tail-bump extent allocator — page 0
|
||
the file header, pages 1-2 the watermark double buffer, pages 3.. data. D6.1
|
||
called for a region table, but the region count is dynamic and unbounded
|
||
(one node arena per index, one slab per collection), and N contiguous
|
||
regions cannot all grow at the tail. With extents there is exactly one
|
||
growth path, therefore exactly one place where the write-then-extend
|
||
discipline lives, and `ls`/`du` stay honest for D6.6's backup story.
|
||
|
||
| # | Commit | Notes |
|
||
|---|---|---|
|
||
| 0 | amend this document | Amendments A1-A4 |
|
||
| 1 | spec-test runner + scorecard | original step 1, still fully independent |
|
||
| 2 | tighten `reserve_for`'s node bound | `n*(depth+2+n/8)+4` demands ~528 MiB of headroom for a 1000-entry multikey batch; free as ArrayList capacity, real file growth once the arena is file-backed |
|
||
| 3 | accessors for node pages and overflow bytes | pure refactor over the existing ArrayList; add the missing comptime asserts pinning `@sizeOf(Node) == 4096` and `@sizeOf(Slot)` |
|
||
| 4 | heap-allocate `Index`; move its scratch buffers out | kills the 5 KB struct memmove in `Collection.remove_index` that dangles a live `Plan.index` |
|
||
| 5 | `_id_` becomes `unique`; add an exact point lookup | Amendment A4, including the replay warning and the test-corpus updates |
|
||
| 6 | leaf payload becomes the slab offset | Amendment A3; the docs map stays as a shadow cross-check that asserts both agree |
|
||
| 7 | streaming candidates + a reverse leaf iterator | a materializing full scan is ~160 MB at the 20 GB gate; cursors are M1 but *streaming a scan* is required here |
|
||
| 8 | `src/pager.zig`: page allocator, mmap over a fixed reservation | engine-unused; unit tests only |
|
||
| 9 | watermark double buffer, catalog stream, free list | engine-unused |
|
||
| 10 | the doc slab lives in the data file | offsets become absolute; the reservation moves ahead of the log append. **First measurable point: A/B bulk-insert throughput and RSS here, before the rest depends on it** |
|
||
| 11 | node arena and overflow live in the data file | 11a overflow, 11b node pages + the id→page table |
|
||
| 12 | copy-on-write above the stable mark; two-generation free list | inert while the stable mark is 0, so testable in isolation first |
|
||
| 13 | checkpoint + watermark; open from the checkpoint; restore `Engine.seq` | log **not** yet truncated, so full replay stays a live safety net and both paths are exercised |
|
||
| 14 | truncate the log after a durable checkpoint | the safety net comes out in its own small commit |
|
||
| 15 | `compact` becomes a data-file rebuild | `compact_snapshot_coll` deleted; same tmp+rename+retry shape |
|
||
| 16 | drop the docs hashmap | original step 7 |
|
||
| 17 | gates | D7's six items; churn gate per amended D6.2 |
|
||
|
||
M0 is done when D7's six items pass. Two harness fixes are prerequisites for
|
||
the gate rather than the work: `bench-run.sh` copies its report over
|
||
`bench-latest.txt` unconditionally, including after a run that only warned,
|
||
so the baseline being defended can be clobbered; and the four B+tree dev
|
||
harnesses (`spill.zig`, `spill2.zig`, `stress.zig`, `fuzz_split.zig`) are in
|
||
no build step, so `zig build test` will not notice an API break in the one
|
||
place that fuzzes splits and >1 KB keys.
|
||
|
||
---
|
||
|
||
## 6. Deferred designs (grill each at its milestone)
|
||
|
||
- **M1 cursors**: cursor id allocation, idle expiration, batchSize
|
||
semantics, getMore against a lagging/compactable engine, cursor state
|
||
lifecycle across compaction.
|
||
- **M2 aggregation**: stage/expression tiers, which spec-test files are
|
||
the gate, whether $lookup/$unwind/facet make the first cut.
|
||
- **M4 transactions**: snapshot isolation over mmap (COW vs undo), read
|
||
concern snapshot, conflict → TransientTransactionError semantics,
|
||
retryable-writes interplay.
|
||
- **M5 change streams**: resume-token design (log seq), live fan-out,
|
||
durability of the resume point.
|
||
- **C API**: exact surface (open/close, command exec, cursor iteration,
|
||
error reporting) — after the compat milestones, on the D1 seam.
|