Collection's slab of malloc'd 8 MiB segments becomes extents in the data file,
and a document's offset becomes an absolute file offset. That makes `doc_bytes`
base + off instead of a binary search over segment starts, and it removes the
hazard the segment list carried: the mapping's base never moves, so a slice into
it cannot be invalidated by growth. phase8 records a dangling-slab-pointer bug
of exactly the shape this deletes.
`slab_reserve` now runs before the log append and `slab_append` after it, and is
infallible. It had no reservation before because appending to an ArrayList could
only fail on OOM; a file-backed slab can also fail on growth, and failing after
the record is durable would report an error for a write the next open produces
anyway.
The data file is still recreated empty on every open and the log still replays in
full, so nothing durable depends on it yet and open/close semantics are
byte-for-byte what they were. The watermark that turns it into a checkpoint comes
next.
--
One bug, and it is worth reading because the milestone keeps producing this
shape. Collection holds a `*Pager`, and `Engine.open` builds an Engine on the
*stack* and returns it by value -- so every pointer taken during replay dangled
the moment it was moved. It surfaced as SIGBUS, then as a corrupt hashmap in the
*second* engine of an unrelated test: nothing resembling its cause. The pager is
heap-allocated now, as `Index` already was, for the same reason.
--
Measured on one harness, 512 MB / 16 KB docs, before and after:
bulk insert throughput 742.6 MB/s -> 746.7 MB/s
createIndex({k: 1}) 26.8 ms -> 16.2 ms
countDocuments({}) 2.1 ms -> 1.1 ms
findOne({k: 500}) indexed 0.75 ms -> 0.53 ms
find({p: range}).count() 6.6 ms -> 4.1 ms
aggregate $group by k 5.8 ms -> 3.7 ms
insertOne (sequential) 0.20 ms -> 0.20 ms
Every row equal or faster. The regression this commit was scheduled early to
catch -- document bytes now reaching disk uncompressed on top of the LZ4 log --
did not appear at this size; bulk insert is flat and the reads gain from one
contiguous mapping instead of separately allocated segments.
What did *not* improve, stated plainly because it is the milestone's headline
claim: RSS is unchanged, 552 MB against 563 MB for 512 MB of data. It cannot
improve yet. Every open still replays the whole log and rewrites the whole slab,
so every page is touched and resident regardless of where it lives. "RSS =
working set" only becomes measurable once an open loads a checkpoint instead of
rebuilding, and the honest test for it is the multi-GB reopen in big.js, not this
microbenchmark.
End-to-end tests with the official MongoDB Node.js driver
These exercise MultiforaDB from a real driver over TCP: full CRUD, query operators, aggregation, error codes, concurrent clients, crash recovery, and the whole lifecycle including server restarts.
Setup
cd tests/e2e
npm init -y >/dev/null
npm install mongodb
Run
Most suites expect a server running on port 27020:
zig build
zig-out/bin/multiforadb --port 27020 --db /tmp/mfdb-e2e.log --ttl-sweep-secs 1 &
node tests/e2e/e2e.js # CRUD + operators + aggregate + errors (29 checks)
node tests/e2e/e2e2.js concurrent # 8 clients: 4 writers + 4 readers (2 checks)
node tests/e2e/e2e2.js crash-a # write 50 docs, then kill -9 the server
node tests/e2e/e2e2.js crash-b # restart and verify all 50 survived
node tests/e2e/e2e3.js # secondary indexes: unique/sparse/compound (16 checks)
node tests/e2e/e2e4.js # TTL indexes: expiry + rejected specs (15 checks)
e2e4.js needs the server started with --ttl-sweep-secs 1 (the default is
60 seconds); the other suites do not care about the flag.
e2e6.js is the full-lifecycle suite and is self-contained: it spawns its
own server on port 27220 with a fresh log, runs the whole feature surface,
restarts the server twice (graceful SIGTERM, then kill -9 mid-write) and
verifies everything survived:
node tests/e2e/e2e6.js # 73 checks, ~15 s, needs no running server
E2E6_PORT=27300 node tests/e2e/e2e6.js # different port if 27220 is taken
Rebuild with zig build after any change under src/ before restarting the
server: zig build test compiles the test binary only and leaves
zig-out/bin/multiforadb stale, so the suites keep running against the old
rules and report failures that the source no longer explains.
e2e2.js concurrent is safe to repeat against a running server (it drops its
collection first); crash-a/crash-b are two halves of one scenario.
Multi-GB collections: big.js
big.js is a load harness, not a pass/fail suite: it spawns a server, bulk
loads up to ~5 GB, and reports insert throughput, the compaction behavior,
server RSS, per-operation latencies, reopen (replay) time, and kill -9
durability.
node tests/e2e/big.js --quick # 268 MB smoke run
node tests/e2e/big.js --size 5g --doc-size 128k --oid --batch 200 \
--compact-threshold 2g # ~5 GB, 40k docs
Options: --size/--doc-size/--batch (k/m/g suffixes), --oid
(ObjectId _ids — see below), --index <field> (secondary index before
loading), --compact-threshold <bytes> (passed to the server),
--port, --keep (keep the db file).
Measured behavior (all documented in the top-level README):
- Build in ReleaseFast —
zig builddefaults to it; a Debug server is 10-200x slower on every path. - Insert throughput collapses under the default 16 MiB compaction
threshold: every ~16 MB of writes rewrites the whole log with one fsync
per record (O(n²) total). With
--compact-threshold 2gthe rate stays flat (hundreds of MB/s at 128 KB docs in ReleaseFast). Raise the threshold for bulk loads. findOne({_id})is O(1) only for ObjectId_ids. Integer_ids are serialization-ambiguous (int32/int64/double compare equal but hash differently), so the docs-map fast path is skipped and every lookup is a full scan. Use the driver's default ObjectId ids on big collections.- The engine holds everything in RAM: ~1-1.2x the data size at 128 KB docs (more at 16 KB docs, where per-document arena overhead dominates). A 5 GB collection needs roughly 6-7 GB of RAM.
- Reopen of a 5 GB log replays in ~10 s (ReleaseFast); every committed write survives kill -9.
Comparing against real MongoDB: compare.js + compare-run.sh
bash tests/e2e/compare-run.sh [size] [doc-size] # e.g. 1g 16k
Starts mongod (brew install mongodb-community) on :27018 and MultiforaDB
on :27019, runs the same driver workload against each (durable writes:
MultiforaDB fsyncs per command, mongod runs with j: true), measures kill -9
reopen for both, and prints a side-by-side table. compare.js alone runs
one side (see its --help-style header comment).
Iteration-to-iteration comparison: bench-run.sh + concurrent.js
bash tests/e2e/bench-run.sh [size] [doc-size] ["clients..."] # e.g. 1g 16k "1 4 8 16 32"
Runs the main suite (compare-run.sh) plus a concurrent durable-write
comparison (concurrent.js, N clients each doing sequential insertOne
with {w:1, j:true} — the group-commit path under real contention), then
writes a machine-readable, versioned report to
tests/e2e/results/bench-<timestamp>.txt and prints a diff of the
MultiforaDB numbers against the previous run (results/bench-latest.txt).
The report has [main] / [concurrency] / [meta] sections with
name<TAB>value rows; bench-run.sh 1g 16k reproduces the phase8 gate
(see results/phase8.txt).