Amendment A5 and the `[M1.1]`/`[M1.2]` blocks. What they record is a result with two halves, and the value is in keeping both: The mechanism works. 934 MB reclaimed over the update run, occupancy at 1.06-1.26x its live data, and the counters that say so are in `serverStatus` rather than inferred. The ratio did not move. 1.94x delete-heavy and 2.46x update-heavy, identical to the end-of-Stage-2 binary measured with the same harness. `file / live` is a high-water mark because the data file never shrinks, and the mark is set in the first round by the one thing reclamation cannot avoid: a rebuild needs a whole second copy of the live data before the first can be freed. So ~2x is the floor of a rebuild-based design and no threshold reaches it -- rebuilding earlier lowers the garbage term and nothing else, rebuilding later raises it. The plan said in advance what to do if this happened, which was to write it down rather than tune, and to name incremental compaction through a doc-id-to-offset indirection layer as the successor. Recorded, with its cost: a second copy-on-write B+tree per collection, a second random read on point lookup, and it undoes A3. A second lever is named that the plan had not: 52% of the steady-state file is space the database owns and is not using, so returning it to the filesystem is worth more here than reclaiming harder. It needs the file never to shrink below what the fallback generation references, which is its own crash-safety pass. D7.4's 1.65x for the delete line is corrected to 1.94x, and the correction is the harness rather than a regression -- the same 1.94x comes out of the binary that predates any of this work. The old ad-hoc version sampled ids to delete blindly, which re-picks dead ones, so it deleted fewer documents than it inserted and measured a collection that was quietly growing. The update line reproduces D7.4 exactly, 2.46 against 2.47. 200-byte documents reclaim nothing, exactly as forecast, and the forecast being written down beforehand is what makes that a result instead of a disappointment. 3.93x on both binaries. Also noted, because the number invites misreading: at that document size the two index trees are comparable to the documents themselves and the file is already 2.18x before any churn -- index structure, not slab garbage.
End-to-end tests with the official MongoDB Node.js driver
These exercise MultiforaDB from a real driver over TCP: full CRUD, query operators, aggregation, error codes, concurrent clients, crash recovery, and the whole lifecycle including server restarts.
Setup
cd tests/e2e
npm init -y >/dev/null
npm install mongodb
Run
Most suites expect a server running on port 27020:
zig build
zig-out/bin/multiforadb --port 27020 --db /tmp/mfdb-e2e.log --ttl-sweep-secs 1 &
node tests/e2e/e2e.js # CRUD + operators + aggregate + errors (29 checks)
node tests/e2e/e2e2.js concurrent # 8 clients: 4 writers + 4 readers (2 checks)
node tests/e2e/e2e2.js crash-a # write 50 docs, then kill -9 the server
node tests/e2e/e2e2.js crash-b # restart and verify all 50 survived
node tests/e2e/e2e3.js # secondary indexes: unique/sparse/compound (16 checks)
node tests/e2e/e2e4.js # TTL indexes: expiry + rejected specs (15 checks)
e2e4.js needs the server started with --ttl-sweep-secs 1 (the default is
60 seconds); the other suites do not care about the flag.
e2e6.js is the full-lifecycle suite and is self-contained: it spawns its
own server on port 27220 with a fresh log, runs the whole feature surface,
restarts the server twice (graceful SIGTERM, then kill -9 mid-write) and
verifies everything survived:
node tests/e2e/e2e6.js # 73 checks, ~15 s, needs no running server
E2E6_PORT=27300 node tests/e2e/e2e6.js # different port if 27220 is taken
e2e7.js is the cursor suite, self-contained for a different reason: cursor
behaviour is only observable with non-default flags. It spawns three servers in
turn -- default flags for batching/streaming/aggregate, then
--cursor-timeout-ms 800 --cursor-sweep-secs 1 --max-open-cursors 4 for idle
expiry and registry capacity, then a restart on the same database to confirm a
cursor does not survive one. Most of it uses raw runCommand, because the
driver hides cursor.id and that is the thing under test:
node tests/e2e/e2e7.js # 86 checks, needs no running server
E2E7_PORT=27310 node tests/e2e/e2e7.js # different port if 27230 is taken
Rebuild with zig build after any change under src/ before restarting the
server: zig build test compiles the test binary only and leaves
zig-out/bin/multiforadb stale, so the suites keep running against the old
rules and report failures that the source no longer explains.
e2e2.js concurrent is safe to repeat against a running server (it drops its
collection first); crash-a/crash-b are two halves of one scenario.
Multi-GB collections: big.js
big.js is a load harness, not a pass/fail suite: it spawns a server, bulk
loads up to ~5 GB, and reports insert throughput, the compaction behavior,
server RSS, per-operation latencies, reopen (replay) time, and kill -9
durability.
node tests/e2e/big.js --quick # 268 MB smoke run
node tests/e2e/big.js --size 5g --doc-size 128k --oid --batch 200 \
--compact-threshold 2g # ~5 GB, 40k docs
Options: --size/--doc-size/--batch (k/m/g suffixes), --oid
(ObjectId _ids — see below), --index <field> (secondary index before
loading), --compact-threshold <bytes> (passed to the server),
--port, --keep (keep the db file).
Measured behavior (all documented in the top-level README):
- Build in ReleaseFast —
zig builddefaults to it; a Debug server is 10-200x slower on every path. - Insert throughput collapses under the default 16 MiB compaction
threshold: every ~16 MB of writes rewrites the whole log with one fsync
per record (O(n²) total). With
--compact-threshold 2gthe rate stays flat (hundreds of MB/s at 128 KB docs in ReleaseFast). Raise the threshold for bulk loads. findOne({_id})is O(1) only for ObjectId_ids. Integer_ids are serialization-ambiguous (int32/int64/double compare equal but hash differently), so the docs-map fast path is skipped and every lookup is a full scan. Use the driver's default ObjectId ids on big collections.- The engine holds everything in RAM: ~1-1.2x the data size at 128 KB docs (more at 16 KB docs, where per-document arena overhead dominates). A 5 GB collection needs roughly 6-7 GB of RAM.
- Reopen of a 5 GB log replays in ~10 s (ReleaseFast); every committed write survives kill -9.
Comparing against real MongoDB: compare.js + compare-run.sh
bash tests/e2e/compare-run.sh [size] [doc-size] # e.g. 1g 16k
Starts mongod (brew install mongodb-community) on :27018 and MultiforaDB
on :27019, runs the same driver workload against each (durable writes:
MultiforaDB fsyncs per command, mongod runs with j: true), measures kill -9
reopen for both, and prints a side-by-side table. compare.js alone runs
one side (see its --help-style header comment).
Iteration-to-iteration comparison: bench-run.sh + concurrent.js
bash tests/e2e/bench-run.sh [size] [doc-size] ["clients..."] # e.g. 1g 16k "1 4 8 16 32"
Runs the main suite (compare-run.sh) plus a concurrent durable-write
comparison (concurrent.js, N clients each doing sequential insertOne
with {w:1, j:true} — the group-commit path under real contention), then
writes a machine-readable, versioned report to
tests/e2e/results/bench-<timestamp>.txt and prints a diff of the
MultiforaDB numbers against the previous run (results/bench-latest.txt).
The report has [main] / [concurrency] / [meta] sections with
name<TAB>value rows; bench-run.sh 1g 16k reproduces the phase8 gate
(see results/phase8.txt).