Two streams of work land together: they are interleaved in storage.zig
and db.zig and only build as a unit.
Already in the working tree before this session:
- ReleaseFast as the default zig build (Debug was 10-200x slower)
- group commit: one fsync per write command instead of per document
- plan_id returned a pointer to a stack temporary; ReleaseFast read
garbage and silently broke findOne({_id: ObjectId})
- perf suite: big.js, compare.js, compare-run.sh, e2e6.js
Phase 1 performance work:
Record integrity hash CRC32 -> XxHash3. std.hash.Crc32 is the
table-driven byte-at-a-time Crc32IsoHdlc, measured at 408 MB/s against
XxHash3's 31 GB/s: 38us versus 0.5us on a 16 KiB document, which was
about two thirds of the entire bulk-insert cost. The record header
grows from u32 crc to u64 hash (header_len 20 -> 24), a breaking
format change. Bulk insert 260 -> 700 MB/s, reopen 1.1 -> 0.5s.
Compaction fsynced once per live document, because Log.open leaves
defer_sync false and compact never set it. It now issues one sync for
the whole rewrite, before the rename that publishes it.
Compaction triggers on the share of the log that is garbage
(live_docs/dead_docs, maintained at evict_doc, the single point where
a document dies) rather than on bytes appended. A fixed byte count is
wrong in both directions: a 1 GB bulk load holds no garbage at all yet
would compact ~64 times under the 16 MiB default, rewriting 1 GB each
time, while a small collection rewritten in place accumulates garbage
indefinitely without ever reaching the count. Pure inserts now never
compact, and the file stays near 1.25x the live data. Bulk load at the
default threshold: 41.6 -> 702.7 MB/s.
remove() never called maybe_compact, so a delete-heavy workload grew
the log without bound.
e2e6's compaction check required the file to bloat past 30 MiB before
being reclaimed, which encoded the old policy and failed on strictly
better behaviour (ends at 15.8 MB against ~12 MB live, was ~30 MB). It
now asserts the file ends near the live size and peaked well above it,
which does not depend on when the trigger fires. Sampling interval 50
-> 10ms: the operations now finish inside the old window.
Verified: 70 unit tests under both ReleaseFast and ReleaseSafe, e2e
29/16/17/3/2 checks, the crash-a/kill -9/crash-b pair, and e2e6 72/72
across three consecutive runs.
100 lines
4.2 KiB
Markdown
100 lines
4.2 KiB
Markdown
# End-to-end tests with the official MongoDB Node.js driver
|
|
|
|
These exercise mongo-lite from a real driver over TCP: full CRUD, query
|
|
operators, aggregation, error codes, concurrent clients, crash recovery, and
|
|
the whole lifecycle including server restarts.
|
|
|
|
## Setup
|
|
|
|
```sh
|
|
cd tests/e2e
|
|
npm init -y >/dev/null
|
|
npm install mongodb
|
|
```
|
|
|
|
## Run
|
|
|
|
Most suites expect a server running on port 27020:
|
|
|
|
```sh
|
|
zig build
|
|
zig-out/bin/mongo-lite --port 27020 --db /tmp/ml-e2e.log --ttl-sweep-secs 1 &
|
|
|
|
node tests/e2e/e2e.js # CRUD + operators + aggregate + errors (29 checks)
|
|
node tests/e2e/e2e2.js concurrent # 8 clients: 4 writers + 4 readers (2 checks)
|
|
node tests/e2e/e2e2.js crash-a # write 50 docs, then kill -9 the server
|
|
node tests/e2e/e2e2.js crash-b # restart and verify all 50 survived
|
|
node tests/e2e/e2e3.js # secondary indexes: unique/sparse/compound (16 checks)
|
|
node tests/e2e/e2e4.js # TTL indexes: expiry + rejected specs (15 checks)
|
|
```
|
|
|
|
`e2e4.js` needs the server started with `--ttl-sweep-secs 1` (the default is
|
|
60 seconds); the other suites do not care about the flag.
|
|
|
|
`e2e6.js` is the full-lifecycle suite and is self-contained: it spawns its
|
|
own server on port 27220 with a fresh log, runs the whole feature surface,
|
|
restarts the server twice (graceful SIGTERM, then kill -9 mid-write) and
|
|
verifies everything survived:
|
|
|
|
```sh
|
|
node tests/e2e/e2e6.js # 73 checks, ~15 s, needs no running server
|
|
E2E6_PORT=27300 node tests/e2e/e2e6.js # different port if 27220 is taken
|
|
```
|
|
|
|
Rebuild with `zig build` after any change under `src/` before restarting the
|
|
server: `zig build test` compiles the test binary only and leaves
|
|
`zig-out/bin/mongo-lite` stale, so the suites keep running against the old
|
|
rules and report failures that the source no longer explains.
|
|
|
|
`e2e2.js concurrent` is safe to repeat against a running server (it drops its
|
|
collection first); `crash-a`/`crash-b` are two halves of one scenario.
|
|
|
|
## Multi-GB collections: `big.js`
|
|
|
|
`big.js` is a load harness, not a pass/fail suite: it spawns a server, bulk
|
|
loads up to ~5 GB, and reports insert throughput, the compaction behavior,
|
|
server RSS, per-operation latencies, reopen (replay) time, and kill -9
|
|
durability.
|
|
|
|
```sh
|
|
node tests/e2e/big.js --quick # 268 MB smoke run
|
|
node tests/e2e/big.js --size 5g --doc-size 128k --oid --batch 200 \
|
|
--compact-threshold 2g # ~5 GB, 40k docs
|
|
```
|
|
|
|
Options: `--size`/`--doc-size`/`--batch` (k/m/g suffixes), `--oid`
|
|
(ObjectId `_id`s — see below), `--index <field>` (secondary index before
|
|
loading), `--compact-threshold <bytes>` (passed to the server),
|
|
`--port`, `--keep` (keep the db file).
|
|
|
|
Measured behavior (all documented in the top-level README):
|
|
|
|
- **Build in ReleaseFast** — `zig build` defaults to it; a Debug server is
|
|
10-200x slower on every path.
|
|
- **Insert throughput collapses under the default 16 MiB compaction
|
|
threshold**: every ~16 MB of writes rewrites the whole log with one fsync
|
|
per record (O(n²) total). With `--compact-threshold 2g` the rate stays
|
|
flat (hundreds of MB/s at 128 KB docs in ReleaseFast). Raise the
|
|
threshold for bulk loads.
|
|
- **`findOne({_id})` is O(1) only for ObjectId `_id`s.** Integer `_id`s are
|
|
serialization-ambiguous (int32/int64/double compare equal but hash
|
|
differently), so the docs-map fast path is skipped and every lookup is a
|
|
full scan. Use the driver's default ObjectId ids on big collections.
|
|
- **The engine holds everything in RAM**: ~1-1.2x the data size at 128 KB
|
|
docs (more at 16 KB docs, where per-document arena overhead dominates).
|
|
A 5 GB collection needs roughly 6-7 GB of RAM.
|
|
- Reopen of a 5 GB log replays in ~10 s (ReleaseFast); every committed
|
|
write survives kill -9.
|
|
|
|
## Comparing against real MongoDB: `compare.js` + `compare-run.sh`
|
|
|
|
```sh
|
|
bash tests/e2e/compare-run.sh [size] [doc-size] # e.g. 1g 16k
|
|
```
|
|
|
|
Starts mongod (`brew install mongodb-community`) on :27018 and mongo-lite
|
|
on :27019, runs the same driver workload against each (durable writes:
|
|
mongo-lite fsyncs per command, mongod runs with `j: true`), measures kill -9
|
|
reopen for both, and prints a side-by-side table. `compare.js` alone runs
|
|
one side (see its `--help`-style header comment).
|