Documents live as canonical BSON bytes in a segmented per-collection slab (fixed 8 MiB segments keep capacity slack under one segment); the docs map holds flat offsets that stay valid across segment growth, and removed documents leave garbage bytes until compaction rewrites. The per-document ArenaAllocator and its second full Pair-tree copy are gone. The matcher walks the stored bytes directly, skipping by length any field the filter does not name (a new bson byte-walker: element_key, skip_value, read_value with borrowed leaves, get_at, and a borrowed spine parse). The byte matcher is differential-tested against the tree matcher on a corpus and shares its operator logic. Stored documents are never materialized on the scan path or in aggregate $match; $group reads group keys and sums straight off the bytes. Sort, projection, findAndModify, updates and index entry generation use a borrowed spine into the slab (or the byte collector, which also replaced collect_values in build_entries). The compaction threshold now counts uncompressed data volume, since a compressed log would otherwise never trigger. Measured (tests/e2e/results/phase5.txt): server RSS 1979 -> 539 MB (2.4x smaller than MongoDB; phase1 baseline 2.0 GB), range-scan 22.5 -> ~12 ms (parity, best run faster than MongoDB), proj 4.1 -> 3.4 ms, createIndex parity. Verified: unit suite in all three modes with zero leaks, the crash pair, e2e6, and the stress/spill programs.
283 lines
14 KiB
Markdown
283 lines
14 KiB
Markdown
# Remaining performance work
|
|
|
|
Status: **items 1 (B+tree), 2 (ordered `_id` index), 3 (block-framed
|
|
compressed log) and 4 (byte storage) are done** — verified in
|
|
`tests/e2e/results/phase2.txt` through `phase5.txt`: updateMany 17.3 →
|
|
1.7 ms, createIndex 62 → 51 ms, `_id` sort+limit 6.2 → 2.4 ms, `db on
|
|
disk` 1025 → 97 MB, and `server RSS` 1979 → 539 MB with the range scan at
|
|
parity (best run faster than MongoDB). Only item 5 (decompose the global
|
|
lock) remains. Each is sized to be landed and verified on
|
|
its own; the ordering constraints between them are the load-bearing part, so
|
|
read those before picking one up.
|
|
|
|
Current numbers and what they mean are in the README; the recorded baseline
|
|
is `tests/e2e/results/phase1.txt`, reproduced with
|
|
`bash tests/e2e/compare-run.sh 1g 16k`.
|
|
|
|
## Ground rules
|
|
|
|
These held for everything already done and should hold here.
|
|
|
|
- **Measure A/B on one harness. Do not trust the model.** Several predictions
|
|
made while planning this work were wrong in both directions: CRC32 cost
|
|
twice what was estimated, one compaction fix turned out to be numerically
|
|
identical to the code it replaced, and a matcher change that measured 1.3x
|
|
in isolation did not move its benchmark row at all. The cheap way to A/B is
|
|
to flip one line, rebuild, run, flip it back.
|
|
- **Mutation-check any test guarding an invariant.** Break the thing the test
|
|
is supposed to catch and confirm it goes red. Several tests here were
|
|
written, looked reasonable, and only proved to have teeth after this.
|
|
- **The index invariant is absolute** (`src/index.zig` header): an index only
|
|
generates candidates, and the full filter is re-applied afterwards.
|
|
Over-approximating is slow. Under-approximating is a wrong answer.
|
|
- **The database must always open.** A unique index that finds duplicates in
|
|
existing data warns and keeps going; replay never refuses to start over
|
|
recoverable damage.
|
|
|
|
Verification for every item:
|
|
|
|
```sh
|
|
zig build test # unit, ReleaseFast
|
|
zig build test -Doptimize=ReleaseSafe # again with safety checks on
|
|
zig build # NOTE: `zig build test` does not refresh the binary
|
|
node tests/e2e/e2e6.js # 72 checks, self-contained, includes kill -9
|
|
```
|
|
|
|
Anything touching the write path or the log format also needs the crash pair
|
|
(`e2e2.js crash-a`, `kill -9`, restart, `e2e2.js crash-b`), and anything
|
|
touching indexes needs `e2e3.js` and `e2e4.js`.
|
|
|
|
---
|
|
|
|
---
|
|
|
|
## 1. B-tree over the encoded keys — DONE
|
|
|
|
Landed as a B+tree in `src/index.zig`: fixed 4 KiB slotted pages in a flat
|
|
u32-addressed `ArrayListUnmanaged(Node)`, an append-only overflow slab for
|
|
records longer than a quarter page (BSON strings reach 16 MB), leaves
|
|
linked for ordered iteration, bulk bottom-up packing for
|
|
`append_doc_entries` + `finish_bulk`, `remove_doc` regenerating entries
|
|
and removing them with a descent plus a leaf-local edit, and no
|
|
rebalancing on delete (empty leaves are unlinked and dropped; internal
|
|
nodes may carry one child). The flat node array stays the contiguous byte
|
|
range that item 4 can write into a checkpoint. Deletes abandon dead pages
|
|
rather than reusing them, so node memory peaks at the tree's peak size,
|
|
exactly what the old entry array's capacity did.
|
|
|
|
Recorded deltas vs `tests/e2e/results/phase1.txt`: `updateMany` 17.3 →
|
|
1.8 ms (2.8x slower than MongoDB → 3.7x faster), `createIndex` 62.4 →
|
|
50.8 ms. Verified with `zig build test` (ReleaseFast and ReleaseSafe), the
|
|
crash pair, `e2e3.js`/`e2e4.js`/`e2e6.js`, plus a 50k-entry stress (bulk
|
|
build, random inserts, random deletes, full drain) and a spill stress
|
|
(2 KiB keys through splits and internal nodes).
|
|
|
|
**Why.** Index entries live in one sorted array, so inserting an entry
|
|
memmoves the tail. Building an index is fine (entries are appended and sorted
|
|
once) and removal is fine (regenerated and found by binary search), but
|
|
*inserting into a collection that already has an index* is quadratic. It also
|
|
blocks item 2.
|
|
|
|
**What.** Replace `Index.entries` with a B+tree whose nodes live in a flat
|
|
`ArrayListUnmanaged(Node)` addressed by `u32`. Slotted 4 KiB pages; keys
|
|
longer than a quarter of a node spill to an overflow slab (BSON strings reach
|
|
16 MB, so this is not optional). Leaves linked for ordered iteration.
|
|
|
|
The flat `u32`-indexed node array is a deliberate choice over pointer-linked
|
|
nodes: it makes the whole index one contiguous byte range that item 4 can
|
|
write into a checkpoint and read back without rebuilding.
|
|
|
|
**Keep.** `append_doc_entries` + `finish_bulk` become bulk leaf packing
|
|
(sort, fill leaves, build interior levels bottom-up). `remove_doc`'s
|
|
regenerate-and-search approach carries over unchanged — only the lookup
|
|
underneath it changes.
|
|
|
|
**Watch.** Deletion is where B+trees get subtly wrong. Not rebalancing on
|
|
delete (leaving underfull leaves, letting compaction reclaim them) is a
|
|
legitimate simplification and much easier to get right; take it unless
|
|
there is a reason not to.
|
|
|
|
**Tests that must keep passing, unchanged:** the randomized `lookup_range`
|
|
check against a brute-force filter, and the `remove_doc`-versus-scan
|
|
differential. Both already exist and both are mutation-checked.
|
|
|
|
---
|
|
|
|
## 2. Ordered `_id` index — DONE
|
|
|
|
Landed as an implicit `_id_` index on every `Collection` (a normal
|
|
`index.Index` with keys `[_id: 1]`, kept out of the secondary `indexes`
|
|
list so `listIndexes`/`dropIndexes`/`createIndex` and the log format are
|
|
unchanged — no `index_create` record, no double listing). Maintained in
|
|
`upsert` (through the same reserve-then-insert protocol as the
|
|
secondaries) and `evict_doc`; rebuilt after replay by `build_all_indexes`
|
|
alongside the secondaries. `index.plan` now takes it as a separate
|
|
argument, so `{_id: ...}` equality, `$in` and ranges use the tree (the
|
|
old serialization-guarded docs-map fast path — `plan_id`,
|
|
`value_fast_path_safe` and friends — is deleted), and `sort({_id: ...})`
|
|
becomes an index-ordered full scan with an early stop. A full `_id` scan
|
|
cannot miss a document: every doc has an `_id` and the index is not
|
|
sparse, and its keys are canonical (`bson.encode_key` gives int32 1,
|
|
int64 1 and double 1.0 identical bytes).
|
|
|
|
Recorded deltas vs `tests/e2e/results/phase2.txt`: `sort({_id:-1}).limit(20)`
|
|
6.2 → 2.4 ms (2.3x slower than MongoDB → parity). Integer/string `_id`
|
|
point lookups, `$in` and ranges no longer full-scan.
|
|
|
|
**Why.** `sort({_id: ...})` still materializes every candidate, and integer
|
|
`_id`s still fall back to a full collection scan on every `findOne`,
|
|
`updateOne` and `deleteOne`.
|
|
|
|
**What.** Give `Collection` an index over the encoded `_id`, maintained in
|
|
`upsert`, `evict_doc` and `apply_record`. Every document has an `_id` and it
|
|
is not sparse, so entry count equals document count and a full index scan
|
|
cannot miss a document — which is what the sort planner's full-scan plan
|
|
requires.
|
|
|
|
Then delete `value_fast_path_safe` and friends (`src/index.zig`). They exist
|
|
only because `serialize_value` gives `int32 1`, `int64 1` and `double 1.0`
|
|
different bytes despite comparing equal. `bson.encode_key` already gives
|
|
them identical bytes, so the guard is obsolete.
|
|
|
|
**Why it depends on item 1.** This index updates on *every* insert. Against a
|
|
sorted array that is a tail memmove each time — roughly 51 GB of memmove over
|
|
65,536 documents. It only looks acceptable because ObjectIds increase
|
|
monotonically and therefore append at the end; random or descending `_id`s
|
|
would collapse bulk insert, currently the project's best result. Do not land
|
|
this against the array without an explicit, documented "append-ordered `_id`s
|
|
only" caveat.
|
|
|
|
**Watch.** A real `_id_` index may start appearing in `listIndexes` and
|
|
writing an `index_create` record to the log. `e2e3.js` asserts on index
|
|
listings — check it before assuming this is invisible. (This landed without
|
|
either: the index stays out of the secondary list, so the listing, drop and
|
|
log surfaces are untouched; verified with `e2e3.js` unchanged.)
|
|
|
|
---
|
|
|
|
## 3. Block-framed compressed log — DONE
|
|
|
|
Landed in `src/storage.zig`: a 16-byte file header (magic, version, codec,
|
|
block target) plus a sequence of 16-byte-header blocks, each holding the
|
|
pre-existing record framing unchanged (`Engine.apply_record` untouched),
|
|
with the integrity hash covering the stored payload bytes so the
|
|
decompressor only ever sees input already proven intact. ~256 KiB target;
|
|
records never straddle blocks (appends accumulate in memory and the block
|
|
seals when the next record would push it past the target). A short read,
|
|
an impossible length, or a hash mismatch in the final block truncates
|
|
cleanly; a mismatch elsewhere is `error.InvalidLog`. A hand-rolled LZ4
|
|
block codec (~1.7 GB/s measured) with a per-block codec byte falling back
|
|
to raw when compression does not help. `Engine.compact` goes through the
|
|
same `Log` API (deferred sync, one commit) and compresses for free.
|
|
|
|
Recorded deltas vs `tests/e2e/results/phase3.txt`: `db on disk` 1025 →
|
|
97 MB (now smaller than MongoDB's own 104 MB); bulk insert 816 → 722 MB/s
|
|
(the compression cost, accepted per the codec note below); reopen 0.8 s
|
|
unchanged. Verified with `zig build test` in all three modes (new LZ4
|
|
round-trip, corrupt-block and torn-tail tests), the crash pair, e2e6
|
|
(kill -9 mid-write), and two full benchmark runs.
|
|
|
|
**Why.** The largest remaining gap: 1.0 GB on disk against MongoDB's 93 MB,
|
|
because payloads are stored raw. Breaking the format is fine.
|
|
|
|
**What.** A file header plus a sequence of blocks; each block holds the
|
|
existing record framing, so `Engine.apply_record` does not change. Target
|
|
~256 KiB per block, and records never straddle blocks.
|
|
|
|
**The hash covers the compressed bytes, as they sit on disk.** If it covered
|
|
the uncompressed bytes you would have to run the decompressor over
|
|
possibly-corrupt input before you could validate it, leaning on the
|
|
decompressor's error handling for durability. Hashing the stored bytes means
|
|
the decompressor only ever sees input already proven intact.
|
|
|
|
**Torn tail versus interior corruption must stay distinguishable**, exactly as
|
|
today: a short read, an impossible length, or a hash mismatch *in the final
|
|
block* means a crash mid-append — truncate and return cleanly. A hash
|
|
mismatch anywhere else is `error.InvalidLog`.
|
|
|
|
**Codec: write an LZ4 block compressor** (~200 lines). Zig 0.16 ships zstd
|
|
*decompression only*, so zstd needs vendoring; `std.compress.flate` has a
|
|
compressor but deflate level 1 runs slower than the current insert rate, so
|
|
it would make writes worse to make the disk smaller. Keep a codec byte in the
|
|
header so raw stays legal (needed anyway when compression does not help) and
|
|
zstd can be swapped in later.
|
|
|
|
**Watch.** `Engine.compact` goes through the same `Log` API and so gets
|
|
compression for free — but it must still defer syncing and commit once.
|
|
|
|
---
|
|
|
|
## 4. Stop giving every document its own arena — DONE
|
|
|
|
Landed: documents live as canonical BSON bytes in a segmented per-collection
|
|
slab (fixed segments keep capacity slack under one segment; the docs map
|
|
holds flat offsets, stable across growth). The matcher walks the bytes
|
|
directly, skipping by length any field the filter does not name, and is
|
|
differential-tested against the tree matcher on a corpus; `$match` in
|
|
aggregate and the scan path never materialize stored documents. Sort,
|
|
projection, updates, findAndModify and index entry generation use a
|
|
borrowed spine (or the byte collector) into the slab. `bson.Document` keeps
|
|
its arena-backed tree meaning for transient docs; stored docs are
|
|
represented by their bytes.
|
|
|
|
Recorded deltas vs `tests/e2e/results/phase4.txt`: `server RSS` 1979 →
|
|
539 MB (2.4x smaller than MongoDB); `range-scan` 22.5 → ~12 ms (parity;
|
|
best run 11.2 vs 14.0); `proj` 4.1 → 3.4 ms. Verified with `zig build
|
|
test` in all three modes (zero leaks; the byte matcher differential), the
|
|
crash pair, e2e6, and the stress/spill programs.
|
|
|
|
**Why.** Two gaps at once. RSS is 2.0 GB against 1.4 GB for a 1.0 GB dataset
|
|
because each document carries an `ArenaAllocator` and a second full copy of
|
|
its data as a `Pair` tree. And the range-scan gap is *not* the matcher — it is
|
|
walking 65,536 documents that each live in a separate allocation, one pointer
|
|
chase apiece.
|
|
|
|
**What.** Store canonical BSON bytes in a per-collection slab; the docs map
|
|
holds a reference into it. Match against those bytes directly, walking the
|
|
element stream and skipping by length any field the filter does not name — for
|
|
the benchmark document that is ~40 bytes touched instead of 16 KB. Add a
|
|
borrowed parse (spine only, pointing into the stored bytes) for sort,
|
|
projection and update. `emit_docs` with no projection becomes a memcpy.
|
|
|
|
**Depends on item 1** in spirit: index entry keys no longer alias document
|
|
arenas (that was fixed when entries moved to encoded bytes), but `Entry.id`
|
|
still aliases the docs map key, and that aliasing needs to be understood
|
|
before documents start moving between slabs.
|
|
|
|
**Watch.** This is the deepest change of the five and touches every test.
|
|
Land the byte matcher behind a flag and differential-test it against the
|
|
existing matcher on a corpus before switching over. Also replace the
|
|
fabricated `bson.Document{ .arena = undefined, ... }` values scattered through
|
|
`query.zig` and `commands.zig` with an explicit view type — they work today
|
|
only because nothing calls `deinit` on them, and they are exactly what breaks
|
|
when `Document` changes meaning.
|
|
|
|
---
|
|
|
|
## 5. Decompose the global lock
|
|
|
|
**Why.** One reader/writer lock covers the entire engine and is held across
|
|
fsync, compaction and reply construction, so writes cannot use more than one
|
|
core.
|
|
|
|
**What.** A catalog lock over the database and collection maps, plus a lock
|
|
per collection. Ordering is catalog then collection, never the reverse, and
|
|
never two collection locks at once — the only cross-collection operations are
|
|
the TTL sweep and compaction, and both must take collections one at a time.
|
|
Then cross-connection group commit: whichever writer finds no commit in
|
|
flight becomes the leader, seals the buffer, compresses, writes and syncs
|
|
once for everyone waiting.
|
|
|
|
**Before touching anything, enumerate what the current code relies on.**
|
|
`Entry.id` aliases the docs map key. `Engine.dup_index` is engine-global state
|
|
written by one call and read by the next. `Log.scratch` is a single shared
|
|
buffer justified by appends being single-writer. Collections are stored *by
|
|
value* in a hash map, so any insert can move them — that one needs fixing
|
|
first regardless.
|
|
|
|
**A durability guarantee changes here, and it should be a decision, not a
|
|
surprise.** Today the log always describes at least as much as memory —
|
|
strictly stronger than MongoDB. Under leader/follower commit a reader can
|
|
observe a write before its fsync completes, which is ordinary `w:1, j:true`
|
|
semantics. Check what `e2e2.js crash-b` actually asserts before changing it.
|