Files
MultiforaDB/tests/e2e
Aleksey Shakhmatov 570900a6ef storage: byte documents in a per-collection slab (roadmap item 4)
Documents live as canonical BSON bytes in a segmented per-collection slab
(fixed 8 MiB segments keep capacity slack under one segment); the docs map
holds flat offsets that stay valid across segment growth, and removed
documents leave garbage bytes until compaction rewrites. The per-document
ArenaAllocator and its second full Pair-tree copy are gone.

The matcher walks the stored bytes directly, skipping by length any field
the filter does not name (a new bson byte-walker: element_key, skip_value,
read_value with borrowed leaves, get_at, and a borrowed spine parse). The
byte matcher is differential-tested against the tree matcher on a corpus
and shares its operator logic. Stored documents are never materialized on
the scan path or in aggregate $match; $group reads group keys and sums
straight off the bytes. Sort, projection, findAndModify, updates and
index entry generation use a borrowed spine into the slab (or the byte
collector, which also replaced collect_values in build_entries). The
compaction threshold now counts uncompressed data volume, since a
compressed log would otherwise never trigger.

Measured (tests/e2e/results/phase5.txt): server RSS 1979 -> 539 MB (2.4x
smaller than MongoDB; phase1 baseline 2.0 GB), range-scan 22.5 -> ~12 ms
(parity, best run faster than MongoDB), proj 4.1 -> 3.4 ms, createIndex
parity. Verified: unit suite in all three modes with zero leaks, the
crash pair, e2e6, and the stress/spill programs.
2026-08-02 22:15:07 +03:00
..

End-to-end tests with the official MongoDB Node.js driver

These exercise mongo-lite from a real driver over TCP: full CRUD, query operators, aggregation, error codes, concurrent clients, crash recovery, and the whole lifecycle including server restarts.

Setup

cd tests/e2e
npm init -y >/dev/null
npm install mongodb

Run

Most suites expect a server running on port 27020:

zig build
zig-out/bin/mongo-lite --port 27020 --db /tmp/ml-e2e.log --ttl-sweep-secs 1 &

node tests/e2e/e2e.js          # CRUD + operators + aggregate + errors (29 checks)
node tests/e2e/e2e2.js concurrent   # 8 clients: 4 writers + 4 readers (2 checks)
node tests/e2e/e2e2.js crash-a      # write 50 docs, then kill -9 the server
node tests/e2e/e2e2.js crash-b      # restart and verify all 50 survived
node tests/e2e/e2e3.js         # secondary indexes: unique/sparse/compound (16 checks)
node tests/e2e/e2e4.js         # TTL indexes: expiry + rejected specs (15 checks)

e2e4.js needs the server started with --ttl-sweep-secs 1 (the default is 60 seconds); the other suites do not care about the flag.

e2e6.js is the full-lifecycle suite and is self-contained: it spawns its own server on port 27220 with a fresh log, runs the whole feature surface, restarts the server twice (graceful SIGTERM, then kill -9 mid-write) and verifies everything survived:

node tests/e2e/e2e6.js              # 73 checks, ~15 s, needs no running server
E2E6_PORT=27300 node tests/e2e/e2e6.js   # different port if 27220 is taken

Rebuild with zig build after any change under src/ before restarting the server: zig build test compiles the test binary only and leaves zig-out/bin/mongo-lite stale, so the suites keep running against the old rules and report failures that the source no longer explains.

e2e2.js concurrent is safe to repeat against a running server (it drops its collection first); crash-a/crash-b are two halves of one scenario.

Multi-GB collections: big.js

big.js is a load harness, not a pass/fail suite: it spawns a server, bulk loads up to ~5 GB, and reports insert throughput, the compaction behavior, server RSS, per-operation latencies, reopen (replay) time, and kill -9 durability.

node tests/e2e/big.js --quick                      # 268 MB smoke run
node tests/e2e/big.js --size 5g --doc-size 128k --oid --batch 200 \
                       --compact-threshold 2g      # ~5 GB, 40k docs

Options: --size/--doc-size/--batch (k/m/g suffixes), --oid (ObjectId _ids — see below), --index <field> (secondary index before loading), --compact-threshold <bytes> (passed to the server), --port, --keep (keep the db file).

Measured behavior (all documented in the top-level README):

  • Build in ReleaseFastzig build defaults to it; a Debug server is 10-200x slower on every path.
  • Insert throughput collapses under the default 16 MiB compaction threshold: every ~16 MB of writes rewrites the whole log with one fsync per record (O(n²) total). With --compact-threshold 2g the rate stays flat (hundreds of MB/s at 128 KB docs in ReleaseFast). Raise the threshold for bulk loads.
  • findOne({_id}) is O(1) only for ObjectId _ids. Integer _ids are serialization-ambiguous (int32/int64/double compare equal but hash differently), so the docs-map fast path is skipped and every lookup is a full scan. Use the driver's default ObjectId ids on big collections.
  • The engine holds everything in RAM: ~1-1.2x the data size at 128 KB docs (more at 16 KB docs, where per-document arena overhead dominates). A 5 GB collection needs roughly 6-7 GB of RAM.
  • Reopen of a 5 GB log replays in ~10 s (ReleaseFast); every committed write survives kill -9.

Comparing against real MongoDB: compare.js + compare-run.sh

bash tests/e2e/compare-run.sh [size] [doc-size]   # e.g. 1g 16k

Starts mongod (brew install mongodb-community) on :27018 and mongo-lite on :27019, runs the same driver workload against each (durable writes: mongo-lite fsyncs per command, mongod runs with j: true), measures kill -9 reopen for both, and prints a side-by-side table. compare.js alone runs one side (see its --help-style header comment).