Files
MultiforaDB/tests/e2e
Aleksey Shakhmatov b4585106f1 storage: block-framed LZ4-compressed log (roadmap item 3)
The log is now a 16-byte file header (magic, version, codec, block
target) plus a sequence of blocks. Each block keeps the pre-existing
record framing unchanged, so Engine.apply_record does not change; records
never straddle blocks (appends accumulate in memory and the block seals
at ~256 KiB). The block header's integrity hash covers the stored payload
bytes exactly as they sit on disk, so the decompressor only ever sees
input already proven intact. Torn tails stay distinguishable from
interior corruption exactly as before: a short read, an impossible
length, or a hash mismatch in the final block truncates cleanly (later
appends overwrite the garbage); a hash mismatch anywhere else is
error.InvalidLog.

The codec is a hand-rolled LZ4 block compressor/decompressor (~1.7 GB/s
measured) with a per-block codec byte falling back to raw when
compression does not help; the header keeps raw legal so zstd can be
swapped in later. Zig 0.16 ships zstd decompression only, and deflate
would cap writes below the insert rate.

Engine.compact goes through the same Log API (deferred sync, one commit)
and compresses for free; sync() seals the pending block before fsyncing,
so the acknowledged-write durability semantics are unchanged (an
unsealed block holds only unacknowledged batch records).

Measured (tests/e2e/results/phase4.txt): db on disk 1025 -> 97 MB, now
smaller than MongoDB's own compressed files; bulk insert 816 -> 722 MB/s
(the accepted compression cost); reopen unchanged at 0.8 s.

Verified: unit suite in all three optimize modes (new LZ4 round-trip,
corrupt-block, and torn-tail truncation tests), the crash pair, e2e6
(kill -9 mid-write), and two full benchmark runs.
2026-08-02 21:35:22 +03:00
..

End-to-end tests with the official MongoDB Node.js driver

These exercise mongo-lite from a real driver over TCP: full CRUD, query operators, aggregation, error codes, concurrent clients, crash recovery, and the whole lifecycle including server restarts.

Setup

cd tests/e2e
npm init -y >/dev/null
npm install mongodb

Run

Most suites expect a server running on port 27020:

zig build
zig-out/bin/mongo-lite --port 27020 --db /tmp/ml-e2e.log --ttl-sweep-secs 1 &

node tests/e2e/e2e.js          # CRUD + operators + aggregate + errors (29 checks)
node tests/e2e/e2e2.js concurrent   # 8 clients: 4 writers + 4 readers (2 checks)
node tests/e2e/e2e2.js crash-a      # write 50 docs, then kill -9 the server
node tests/e2e/e2e2.js crash-b      # restart and verify all 50 survived
node tests/e2e/e2e3.js         # secondary indexes: unique/sparse/compound (16 checks)
node tests/e2e/e2e4.js         # TTL indexes: expiry + rejected specs (15 checks)

e2e4.js needs the server started with --ttl-sweep-secs 1 (the default is 60 seconds); the other suites do not care about the flag.

e2e6.js is the full-lifecycle suite and is self-contained: it spawns its own server on port 27220 with a fresh log, runs the whole feature surface, restarts the server twice (graceful SIGTERM, then kill -9 mid-write) and verifies everything survived:

node tests/e2e/e2e6.js              # 73 checks, ~15 s, needs no running server
E2E6_PORT=27300 node tests/e2e/e2e6.js   # different port if 27220 is taken

Rebuild with zig build after any change under src/ before restarting the server: zig build test compiles the test binary only and leaves zig-out/bin/mongo-lite stale, so the suites keep running against the old rules and report failures that the source no longer explains.

e2e2.js concurrent is safe to repeat against a running server (it drops its collection first); crash-a/crash-b are two halves of one scenario.

Multi-GB collections: big.js

big.js is a load harness, not a pass/fail suite: it spawns a server, bulk loads up to ~5 GB, and reports insert throughput, the compaction behavior, server RSS, per-operation latencies, reopen (replay) time, and kill -9 durability.

node tests/e2e/big.js --quick                      # 268 MB smoke run
node tests/e2e/big.js --size 5g --doc-size 128k --oid --batch 200 \
                       --compact-threshold 2g      # ~5 GB, 40k docs

Options: --size/--doc-size/--batch (k/m/g suffixes), --oid (ObjectId _ids — see below), --index <field> (secondary index before loading), --compact-threshold <bytes> (passed to the server), --port, --keep (keep the db file).

Measured behavior (all documented in the top-level README):

  • Build in ReleaseFastzig build defaults to it; a Debug server is 10-200x slower on every path.
  • Insert throughput collapses under the default 16 MiB compaction threshold: every ~16 MB of writes rewrites the whole log with one fsync per record (O(n²) total). With --compact-threshold 2g the rate stays flat (hundreds of MB/s at 128 KB docs in ReleaseFast). Raise the threshold for bulk loads.
  • findOne({_id}) is O(1) only for ObjectId _ids. Integer _ids are serialization-ambiguous (int32/int64/double compare equal but hash differently), so the docs-map fast path is skipped and every lookup is a full scan. Use the driver's default ObjectId ids on big collections.
  • The engine holds everything in RAM: ~1-1.2x the data size at 128 KB docs (more at 16 KB docs, where per-document arena overhead dominates). A 5 GB collection needs roughly 6-7 GB of RAM.
  • Reopen of a 5 GB log replays in ~10 s (ReleaseFast); every committed write survives kill -9.

Comparing against real MongoDB: compare.js + compare-run.sh

bash tests/e2e/compare-run.sh [size] [doc-size]   # e.g. 1g 16k

Starts mongod (brew install mongodb-community) on :27018 and mongo-lite on :27019, runs the same driver workload against each (durable writes: mongo-lite fsyncs per command, mongod runs with j: true), measures kill -9 reopen for both, and prints a side-by-side table. compare.js alone runs one side (see its --help-style header comment).