Files
MultiforaDB/tests/e2e/results/m0-gates.txt
A.Shakhmatov a748a3d08c db/pager/tests: cleanup pass over the free list
No behaviour is meant to change and the gate confirms it: 1.94x / 679.2 MB
reclaimed / 9 rebuilds and 2.46x / 934.0 MB / 13 on the two 16 KiB lines,
0.0 MB on the 200-byte line, all identical to the numbers recorded for them.

Deduplication. `pages_for` was written in db.zig and again in pager.zig and
twice more inline; there is now one, public, and the two pre-existing copies
call it. `pages_per_map_align` replaces three hand-rolled `map_align /
page_size`. `SlabRun.window_first` was a stored field that could never legally
disagree with `first` and was maintained by hand at two sites -- now a method.
`keep_piece` re-derived `SlabRun.window_count` character for character; it
calls it. `insert_run` scanned linearly for a position `run_of` binary-searches
for, which made loading a fragmented catalog quadratic; both now go through one
`run_lower_bound`. Freeing a run's window map was written four times; one
helper. The 20% rebuild share was stated in `note_compact` and again in
`wants_rebuild`, with a comment arguing at length that they must be the same
number -- `worth_rewriting` makes that structural.

Efficiency. The identity assert in `reclaim_windows` called `dead_located()`,
an O(every window) walk, and `assert_msg` is live in ReleaseFast -- so it
doubled the scan the reclamation was about to make (2.75 MB streamed twice per
reclaiming checkpoint at the 21 GB the gate targets). `dead_located` is now a
maintained counter, the check is O(1) in every build, and the scan cross-checks
it while it is there. `SlabRun.full` lets a run with nothing to give be copied
without its counters being read at all, so the common case is O(runs) rather
than O(windows).

The pager's two allocation policies were hand-copying the claim step, and the
copy had already lost two of the three preconditions -- `alloc_slab_run` never
checked `pages <= reserved_pages`. Both now go through `claim_locked`.

`reclaimed_bytes` moves from Collection to Engine, beside `compactions`, which
is how it is read and the only place it can be honest: a life-of-the-process
total must not lose a dropped collection's share. Both join `Counters`, so
`slab_stats` stops opening `counter_lock` by hand.

The ownership assertion in `write_catalog` was gated on `is_test or Debug`, a
predicate nothing else in the codebase uses, which left the one silent failure
this design can produce unchecked in ReleaseSafe. It is now `!= ReleaseFast`,
the line `protect_stable` already draws. Measured: no change to the suite's
runtime.

Altitude. `note_checkpoint` was called from exactly one place, the tail of
`upsert` -- so a delete armed no checkpoint by any route, which is why
reclamation only ever ran when the *rebuild* trigger fired and the rebuild then
reset the window map it would have used. `remove` and the TTL sweep arm one
now, next to the `note_compact` calls that were added for the same omission a
milestone ago. `compact`'s leading checkpoint stays, demoted in its comment
from the mechanism to the local ordering it actually guarantees.

serverStatus reports `allocTailBytes`/`freeReadyBytes` instead of page counts,
so the harness stops hard-coding 4096 -- the kind of constant this milestone
was blindsided by once already.

tests/e2e/churn.js: `deleteMany({_id: {$in: [5000 ids]}})` exceeded
`index.max_combos`, so the planner refused the index and every delete became a
full collection scan re-filtering each document against 5000 members. That was
the entire runtime of the harness. One delete spec per id instead: the 40k x
16 KiB gate goes 57 s -> 6 s, and the 150k x 200 B line 483 s -> 3 s, with
identical output. Also: the per-round `countDocuments` is gone (the harness
knows the count), and the server log is a bounded ring rather than a rope that
grows with everything the server ever said.

Reverted from the review: reusing one MongoClient across the startup poll. A
client whose first connect fails tears its topology down and every later
command on it fails identically, so it turns "not up yet" into "never comes
up" -- it broke the first run. The reason is now a comment.

187/187 unit tests in ReleaseFast and ReleaseSafe, 83/83 fuzz, e2e 49, e2e2
concurrent 2 and the crash pair, e2e3 16, e2e4 17, e2e6 72, e2e7 86,
crash-fuzz 60 cycles.
2026-08-09 19:08:15 +03:00

258 lines
15 KiB
Plaintext

# M0 gate results — mmap + WAL storage foundation (PLAN D7)
#
# Machine: Apple Silicon, macOS 25.5.0, 16 KiB system pages, APFS.
# Server: MultiforaDB at the commit named below, ReleaseFast.
# Driver: mongodb@7.5.0 (pinned in tests/e2e/package-lock.json).
# mongod: 8.3.7, for the parity rows.
#
# Read this alongside PLAN.md §5. Every number here is reproducible with the
# command printed under it; where a gate was not met, the number is recorded as
# measured and the reason is stated rather than the workload being tuned until
# it passed.
[D7.1] unit tests, ReleaseFast and ReleaseSafe
zig build test -Doptimize=ReleaseFast 122/122 pass
zig build test -Doptimize=ReleaseSafe 122/122 pass
zig build fuzz fuzz_split, spill, spill2, stress clean
Both modes matter: `protect_stable` (the mprotect belt over the published
image) is comptime-off in ReleaseFast, and `std.posix.mprotect` not existing
in Zig 0.16 was a ReleaseSafe-only compile error.
[D7.2] end-to-end suites — all green
reproduce: zig-out/bin/multiforadb --port 27020 --db /tmp/mfdb-e2e.log \
--ttl-sweep-secs 1 & then node tests/e2e/<suite>
e2e.js CRUD, operators, aggregate, errors 35/35
e2e3.js secondary indexes: unique/sparse/compound 16/16
e2e4.js TTL indexes: expiry and rejected specs 17/17
e2e5.js miscellaneous command surface 3/3
e2e2.js concurrent, 4 writers + 4 readers 2/2
e2e2.js crash-a / crash-b (kill -9, then verify) 1/1 then 3/3
e2e6.js compaction and log growth (own server) 72/72
[D7.3] large-collection smoke — big.js
reproduce: node tests/e2e/big.js --size 20g --doc-size 16k
20g run 4g run
documents 1,310,720 262,144 (x 16 KiB)
live data 21.47 GB 4.29 GB
data file 21.75 GB 4.35 GB (+1.3%, +1.4%)
log file, min -> final 16 B -> 2.5 MB 16 B -> 7.7 MB
bulk insert 411.0 MB/s 560.3 MB/s
peak server RSS during load 18,492 MB 3,318 MB
kill -9 then reopen 0.5 s 0.5 s
RSS after reopen 237 MB (1.1%) 130 MB (3.0%)
count after restart 1,310,720 OK 262,144 OK
last doc intact after restart payload 16254B OK
kill -9 after 200 acked writes 200/200 200/200
This is the milestone's central claim, measured: an open costs the working
set, not the size of the database, and it no longer replays the whole log.
Reopen is flat at 0.5 s from 4 GB to 21.5 GB, and RSS after reopen falls as a
*share* of the data as the database grows — which is what "RSS = working set"
has to mean to be worth anything.
Before M0 the same shape of run reported 523 MB resident after reopen for a
512 MB database, because recovering each document's `_id` meant reading every
document at open. That scan is what the mmap foundation deleted.
Two other things the 20g run shows, honestly. Peak RSS during the *load* is
18.5 GB: writing 21 GB touches 21 GB of pages and the kernel keeps them until
it needs the memory, so a bulk load is not where the mmap win shows up. And a
cold full scan of 21 GB costs ~55 s — countDocuments, the range scan and the
unindexed updateMany all land there, ~390 MB/s off the disk — which is correct
and is the reason M1's cursors and index-only paths matter.
[D7.4] churn gate — doc/slab garbage (PLAN amendment A2 retargets D6.2 here)
40,000 x 16 KiB documents (655 MB live), one secondary index, steady state
measured as data-file size / live data after the ratio stopped moving.
before the fixes after
delete half and refill, 6 rounds 4.10x, climbing 1.65x, flat
random $set over 5x the collection 3.58x, climbing 2.47x, flat
"Climbing" is the whole finding: nothing was being reclaimed at all. Three
bugs, all fixed in `db/pager: reclaim what churn abandons` — a compaction
trigger that had read `log.data_bytes` since before the log was truncated at
every checkpoint, a numeric stable mark that made recycled pages look
published, and a first-fit free list that let one-page requests dismantle
the extents. A fourth fix (a rebuild publishes twice, so the space it frees
is reusable immediately) took the interleaved case from 3.58x to 2.47x.
The gate hoped for ~1.3x and this is above it, structurally: a rebuild needs
a whole second copy of the live data before the first can be freed, and
two-generation retention holds the old copy through two more publishes. The
gate's purpose was to decide whether doc-level free lists are needed after
M0. They are; that is an M1 item. What M0 owes is a bound, and there is one.
[D7.5] benchmark parity vs the pre-mmap engine
reproduce: bash tests/e2e/bench-run.sh 1g 16k "1 4 8 16 32"
baseline: tests/e2e/results/phase8.txt (c8d547f, all-in-RAM engine)
M0: three runs, tests/e2e/results/bench-2026080{3-223947,3-225128,3-225322}.txt
Read the range, not a single delta. Two consecutive runs of the *same* binary
moved the sub-10 ms rows by 27-51% on this machine, so any one comparison
reads whatever the noise did that minute. Best-of-three against the baseline:
row pre-mmap M0 min..max best
insertOne (sequential) x200 0.20 ms 0.20..0.22 ms +0%
bulk insert throughput 732.4 MB/s 463.4..555.5 MB/s -24%
createIndex({k: 1}) 74.4 ms 27.9..28.9 ms -62%
countDocuments({}) 3.1 ms 2.6..3.6 ms -16%
findOne({_id: <ObjectId>}) 0.59 ms 0.60..0.80 ms +2%
findOne({k: 500}) (indexed) 0.56 ms 0.54..0.73 ms -4%
find({p: {$gte,$lt}}).count() (scan) 13.0 ms 11.2..14.5 ms -14%
find({}).sort({_id:-1}).limit(20) 2.2 ms 1.5..1.9 ms -32%
find({}, {proj}).limit(1000) 3.5 ms 3.6..3.8 ms +3%
aggregate $group by k 8.3 ms 7.4..11.2 ms -11%
updateOne({_id}) x50 0.14 ms 0.15..0.18 ms +7%
updateMany({k: 7}, {$inc}) 1.9 ms 1.9..2.0 ms +0%
deleteOne({_id}) + insertOne 0.58 ms 0.63..0.69 ms +9%
node client RSS 158 MB 150..163 MB -5%
concurrent durable insertOne, docs/s:
clients 1 8,435 7,709..8,373 -1%
clients 4 22,480 20,122..20,880 -7%
clients 8 25,920 26,210..26,349 +2%
clients 16 31,079 27,794..29,594 -5%
clients 32 34,041 32,235..32,817 -4%
One reproducible regression: bulk insert, -24%, steady across all three runs.
This is PLAN risk 1 exactly as written -- document bytes now reach the disk
uncompressed in the data file on top of the LZ4 log, and that writeback
bandwidth is new. Every other row is inside the run-to-run spread. The gate
as stated was "no phase8 row regressed"; it is met for the read and latency
rows and not for bulk load, which is the trade the milestone makes and which
risk 1 anticipated. Untried mitigations, in the order worth trying:
MADV_HUGEPAGE / a larger growth chunk, and not writing doc bytes twice
(log and slab) for a bulk path that could log an extent reference instead.
createIndex at -62% is the other reproducible number, and it comes from the
same change: a bulk build now packs pages in the mapping instead of growing
an ArrayList.
[D7.6] MongoDB spec-test scorecard
reproduce: bash tests/spec/fetch.sh && node tests/spec/run.js --scorecard
total 131 pass 161 fail 195 skip 175 files 0 errored
Byte-identical to the scorecard recorded before the storage rewrite. That is
the intended result and it is worth stating plainly: M0 replaced the document
store, the node arena and the overflow slab, added copy-on-write, a watermark
and a checkpoint, and moved every index leaf's payload -- and changed no
observable CRUD or aggregate semantics.
# WHAT THE GATES FOUND
# Five bugs, none of which any unit test or e2e suite had reached:
# 1. The compaction trigger had been dead since commit 14 (log truncation
# zeroes the counter it gated on). Nothing reclaimed doc-slab garbage.
# 2. `page_mut_cow` asked `p >= stable_pages`, which is false for a recycled
# page, so every write to one copied and freed it again.
# 3. First fit let copy-on-write's one-page requests shave the extents the
# doc slab needs, so the free list drained every generation and the file
# grew anyway.
# 4. A rebuild's freed space took two unrelated checkpoints to become
# reusable, so the next rebuild grew the file instead of reusing it.
# 5. `reserve_pages`' promise was a single counter on the pager. Two upserts
# on different collections release it independently, so the first to
# finish revoked the second's promise mid-write -- aborting the server at
# four concurrent clients, on the first concurrent benchmark since the
# data file landed. PLAN risk 3, whose mitigation was never built.
# Four of the five are invisible without a workload that runs long enough to
# reach a steady state, or wide enough to have two writers. That is the
# argument for keeping both the churn gate and the concurrent benchmark in the
# gate list rather than treating them as optional.
#
# ONE COMPATIBILITY GAP FOUND HERE, FIXED AFTER THE GATE
# `update.apply` rejected any update document whose first key was not `$`, so
# replacement-style writes -- replaceOne, findOneAndReplace, bulkWrite's
# replaceOne -- failed with "bad update". Found while building the churn
# workload, out of scope for a storage milestone, and fixed immediately after in
# `update: replacement-style writes` and `db: a write that changes nothing is
# not a write`. The scorecard above is the M0 figure and is left as measured;
# those two commits took it to 163 pass / 129 fail, and `tests/spec/scorecard.txt`
# always holds the current one.
# ===========================================================================
# M1 — doc-level free list (PLAN amendment A5)
# ===========================================================================
#
# Same machine, same driver. Server at the M1 commit named per block,
# ReleaseFast. These are the numbers D7.4 said an M1 item owed.
[M1.1] churn gate — the doc-level free list
reproduce: node tests/e2e/churn.js --docs 40000 --doc-size 16k --index \
--mode delete-refill --rounds 6
node tests/e2e/churn.js --docs 40000 --doc-size 16k --index \
--mode update --multiple 5
baseline: the same harness against the end-of-Stage-2 binary (e416ad1),
which has no reclamation and no `multifora` section.
The harness is committed this time (`tests/e2e/churn.js`), which is half the
point of the block: D7.4's numbers were real and unrepeatable. Each line
above runs in 6-10 s.
Stage 2 M1 target
delete half and refill, 6x 1.94x 1.94x <= 1.45x NOT MET
random $set over 5x the coll. 2.46x 2.46x <= 1.60x NOT MET
Both flat, drift +0.00x over the last three rounds.
D7.4 recorded 1.65x for the delete line. This harness reads 1.94x for the
*same binary* D7.4 was measured against the descendants of, so that gap is
the harness, not a regression: the ad-hoc version sampled ids to delete
blindly, which re-picks already-dead ids, deletes fewer than it inserts and
measures a collection that is quietly growing. The update line reproduces
D7.4 exactly (2.46 vs 2.47).
THE RATIO DID NOT MOVE AND THE MECHANISM WORKS. Both are true, and the
counters are what separate them:
round 6, update line: reclaimed 934.0MB dead 83.9MB
allocTail 1523.5MB freeReady 797.0MB
inUse 1.16x file/live 2.46x
Reclamation returned 934 MB over the run and the collection is occupying
1.16x its live data. What 2.46x measures is the data file's high-water mark,
and the file never shrinks. The mark is set once, in round 1, by the one
thing reclamation cannot avoid: a rebuild needs a whole second copy of the
live data before the first copy can be freed. 626 MB live + the garbage
standing at the moment it fires + 626 MB of copy is the number, and it is
reached before any free pool exists to build the copy out of.
So the floor for a rebuild-based design is ~2x, and no threshold reaches it.
Rebuilding earlier lowers the garbage term and raises nothing; rebuilding
later raises it. The plan anticipated this exact outcome and said what to do
about it, which is to write it down rather than tune: the remaining lever is
incremental compaction -- a doc-id-to-offset indirection layer, so a rebuild
moves documents without a second copy of everything. That is amendment A5's
successor and it is a milestone of its own, not a knob.
A second lever, cheaper and not attempted: give free space back to the
filesystem. `freeReady` stands at 797 MB with `allocTail` flat, so 52% of the
file is space the database owns and is not using. Returning the tail-adjacent
part of it needs the file never to shrink below what the fallback generation
references, which is a crash-safety argument and its own design pass.
What did change, and is the reason the mechanism is worth keeping:
delete-refill, 12k x 16 KiB reclaimed 1 MB -> 256 MB (six rounds)
The first figure is what reclamation achieved before `compact` was made to
checkpoint before it copies. A checkpoint is what reclaims and a checkpoint
is armed by log volume; a delete logs only an `_id`, so deleting half a
collection moved the log by a couple of megabytes, no checkpoint ran, and the
rebuild got there first every time and reset the window map it would have
used. The harness found that on its first serious run, which is the argument
for committing it.
[M1.2] churn gate — 200-byte documents
reproduce: node tests/e2e/churn.js --docs 150000 --doc-size 200 --index \
--mode delete-refill --rounds 4 (3 s)
Predicted in advance, in the plan, as a pass rather than a fault:
reclaimed 0.0 MB over four rounds, exactly as forecast.
ratio 3.93x on both the Stage 2 binary and M1 -- identical, flat.
Reclamation hands back whole system pages. A 16 KiB page on this machine
holds ~70 documents of 200 bytes and the chance that all 70 are dead at once
under uniform deletion is nil, so nothing is ever handed back. The forecast
said the ratio would not improve and the counters would show a mechanism
that correctly does nothing, rather than one that silently misfires; that is
what they show.
Read the ratio on this line with care. `live` counts document bytes, and at
200 bytes the two index trees are comparable in size to the documents
themselves -- the file is already 2.18x at load, before any churn. That
overhead is index structure, not slab garbage, and it is not what this gate
is about.
The payoff of window reclamation scales as doc_size / map_align, so a 4 KiB
system page (x86-64 Linux) reads four times better on the same code. Every
number in this file is Apple Silicon with 16 KiB pages.