M1: doc-level free list, sessions, and a spec runner that no longer overstates #1
99
PLAN.md
99
PLAN.md
@@ -320,6 +320,71 @@ takes the first one's place. That commit is where this needs handling — a
|
|||||||
pre-flight scan for compare-equal `_id`s, refusing to drop the map silently
|
pre-flight scan for compare-equal `_id`s, refusing to drop the map silently
|
||||||
while any exist — not here.
|
while any exist — not here.
|
||||||
|
|
||||||
|
### Amendment A5 — the doc-level free list, and what it did not fix (amends A2, closes D7.4)
|
||||||
|
|
||||||
|
D7.4 left M0 with a bound rather than a target: 1.65× delete-heavy, 2.47×
|
||||||
|
update-heavy against a hoped-for ~1.3×, and the stated conclusion that
|
||||||
|
doc-level free lists were an M1 item. They are built. The mechanism is
|
||||||
|
measured, it works, and **the steady-state ratio did not move**. Both halves
|
||||||
|
of that are the amendment.
|
||||||
|
|
||||||
|
**What was built.** A collection's slab carries a dense map of dead bytes per
|
||||||
|
`map_align` window — two bytes per window, so 2.7 MB for a 21 GB slab — and a
|
||||||
|
checkpoint hands back every window with nothing live left in it, splitting the
|
||||||
|
runs around what is kept. The window is the unit because it is the smallest
|
||||||
|
thing that can be given back at all: `mark_appendable` refuses an unaligned
|
||||||
|
start and `protect_stable` rounds outwards. Counting is the whole liveness
|
||||||
|
test, because `evict_doc` removes a document's index entries before marking its
|
||||||
|
bytes dead, so "no live bytes in this window" and "nothing references these
|
||||||
|
bytes" are the same statement. Reclamation lives inside `checkpoint` rather
|
||||||
|
than beside it so that the run split and the `free_pages` become durable under
|
||||||
|
one `publish`; there is no new record type, no new catalog version and no
|
||||||
|
replay path. `alloc_slab_run` is a second policy in the same allocator, because
|
||||||
|
`take_free`'s best fit — which exists to stop one-page copy-on-write requests
|
||||||
|
dismantling the extents — can never match a request for 2048 pages against
|
||||||
|
runs that come back a few windows at a time.
|
||||||
|
|
||||||
|
**What it does not fix, and why no threshold reaches it.** The data file never
|
||||||
|
shrinks, so `file / live` is a high-water mark, and the mark is set once by the
|
||||||
|
one thing reclamation cannot avoid: a rebuild needs a whole second copy of the
|
||||||
|
live data before the first can be freed. Live + garbage-at-trigger + copy is
|
||||||
|
the peak, and it is reached in the first round, before any free pool exists to
|
||||||
|
build the copy out of. Rebuilding earlier lowers the garbage term and nothing
|
||||||
|
else; rebuilding later raises it. So ~2× is the floor of a rebuild-based
|
||||||
|
design, and tuning is the wrong instrument. Measured occupancy tells the other
|
||||||
|
half of the story: 1.06–1.26× in use against a 2.46× file, with 934 MB
|
||||||
|
reclaimed over the run.
|
||||||
|
|
||||||
|
**The successor, named here so the next session does not re-derive it.**
|
||||||
|
Incremental compaction through a doc-id → offset indirection layer, which is
|
||||||
|
rejected option (b) of the M1 design, promoted. It is the only thing that
|
||||||
|
removes the second copy: a rebuild becomes a move of one document at a time
|
||||||
|
with the map updated behind it. The cost is the one that got it rejected — the
|
||||||
|
map has to be persistent and crash-safe, i.e. a second copy-on-write B+tree per
|
||||||
|
collection and a second random read on the point-lookup path — and it undoes
|
||||||
|
A3. That is a milestone, not a knob. Second and cheaper: return free space to
|
||||||
|
the filesystem, since 52% of the steady-state file is space the database owns
|
||||||
|
and is not using; it needs the file never to shrink below what the fallback
|
||||||
|
generation references, which is its own crash-safety design pass.
|
||||||
|
|
||||||
|
**Small documents behave exactly as forecast**, and the forecast being written
|
||||||
|
down in advance is what makes it a result. 200-byte documents reclaim nothing
|
||||||
|
at all — a 16 KiB system page holds ~70 of them and they never all die at once
|
||||||
|
— and the counters show a mechanism correctly doing nothing rather than one
|
||||||
|
misfiring. The payoff scales as `doc_size / map_align`, so 4 KiB pages read
|
||||||
|
four times better on the same code.
|
||||||
|
|
||||||
|
**One thing the gate found that the design had not.** A checkpoint is what
|
||||||
|
reclaims and a checkpoint is armed by log volume, but a delete logs only an
|
||||||
|
`_id`. Deleting half a 190 MB collection moved the log by a couple of megabytes
|
||||||
|
so no checkpoint ran, the garbage sailed past the rebuild threshold, and the
|
||||||
|
rebuild reset the window map it would have used — six rounds, six rebuilds,
|
||||||
|
1 MB reclaimed. `compact` now checkpoints before it copies, which is also the
|
||||||
|
right order on its own terms: the cheap half of the job first, and the
|
||||||
|
per-collection gate judges what reclamation left. Same six rounds, 256 MB.
|
||||||
|
Numbers and reproduction in `tests/e2e/results/m0-gates.txt` under `[M1.1]` and
|
||||||
|
`[M1.2]`.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## 3. Milestones and gates
|
## 3. Milestones and gates
|
||||||
@@ -721,6 +786,16 @@ between `compact` and `checkpoint`; it is out of this scope because it wants
|
|||||||
its own design pass, and because the free list must not add a second instance
|
its own design pass, and because the free list must not add a second instance
|
||||||
of the same shape.
|
of the same shape.
|
||||||
|
|
||||||
|
*Where this stands after the free list.* It did add a second instance —
|
||||||
|
reclamation frees pages as a checkpoint phase, and two checkpoints can be in
|
||||||
|
flight — so that half is closed: `checkpoint` takes a lock of its own. Two
|
||||||
|
things came out of doing it. The publish was never the exposure, because it
|
||||||
|
already runs under `log_lock`; and the whole class is now *detectable* rather
|
||||||
|
than only arguable, because `write_catalog` asserts per run that the pager has
|
||||||
|
not already been given it, in test and Debug builds. That assertion is proven
|
||||||
|
to fire. The original instance — `compact`'s rebuild walk against a concurrent
|
||||||
|
checkpoint — is unchanged and still wants the design pass.
|
||||||
|
|
||||||
### The spec runner starts reading `expectEvents`
|
### The spec runner starts reading `expectEvents`
|
||||||
|
|
||||||
354 of the 487 cases declare `expectEvents` and the runner read none of them,
|
354 of the 487 cases declare `expectEvents` and the runner read none of them,
|
||||||
@@ -823,15 +898,21 @@ has to be its own commit with its own re-recorded scorecard.
|
|||||||
the anchor rewritten — resuming at it returned updated documents twice,
|
the anchor rewritten — resuming at it returned updated documents twice,
|
||||||
caught by draining a collection being updated underneath.
|
caught by draining a collection being updated underneath.
|
||||||
|
|
||||||
Still open in M1: the doc-level free list. The eight reclamation bugs above
|
The doc-level free list is built; see amendment A5 for what it did and did
|
||||||
were cleared first, as preconditions for the free list rather than as work of
|
not achieve, and `[M1.1]`/`[M1.2]` in the results file for the numbers. The
|
||||||
their own; command-monitoring (`expectEvents`) landed next, so that what
|
eight reclamation bugs above were cleared first, as preconditions for it
|
||||||
followed is measured by an instrument no longer known to overstate.
|
rather than as work of their own; command-monitoring (`expectEvents`) landed
|
||||||
**A prerequisite the free list must honour**, recorded here while it is
|
next, so that what followed is measured by an instrument no longer known to
|
||||||
still being designed: *an offset that was ever a record start must remain a
|
overstate.
|
||||||
record start.* `doc_bytes` reads a `u32` length prefix in place, so an
|
**The prerequisite it had to honour** — *an offset that was ever a record
|
||||||
offset landing mid-record after a re-split is a garbage-length read rather
|
start must remain a record start*, because `doc_bytes` reads a `u32` length
|
||||||
than a wrong answer — and an offsets cursor holds exactly such offsets.
|
prefix in place and an offsets cursor holds exactly such offsets — is met
|
||||||
|
structurally rather than by checking: a window is handed back only when every
|
||||||
|
byte in it is dead, which means every document touching it has already been
|
||||||
|
through `evict_doc` and out of every index. What remains is the cursor
|
||||||
|
holding a *saved* offset list, and that is answered the way a rebuild answers
|
||||||
|
it, by bumping `layout_epoch` when and only when a collection actually gave
|
||||||
|
something back.
|
||||||
- **M1 sessions** — *settled and implemented.* `lsid` is parsed, validated and
|
- **M1 sessions** — *settled and implemented.* `lsid` is parsed, validated and
|
||||||
deliberately acted on in no way; `txnNumber`, `startTransaction` and
|
deliberately acted on in no way; `txnNumber`, `startTransaction` and
|
||||||
`autocommit` are refused; `endSessions` validates the array it discards.
|
`autocommit` are refused; `endSessions` validates the array it discards.
|
||||||
|
|||||||
@@ -157,3 +157,100 @@
|
|||||||
# not a write`. The scorecard above is the M0 figure and is left as measured;
|
# not a write`. The scorecard above is the M0 figure and is left as measured;
|
||||||
# those two commits took it to 163 pass / 129 fail, and `tests/spec/scorecard.txt`
|
# those two commits took it to 163 pass / 129 fail, and `tests/spec/scorecard.txt`
|
||||||
# always holds the current one.
|
# always holds the current one.
|
||||||
|
|
||||||
|
|
||||||
|
# ===========================================================================
|
||||||
|
# M1 — doc-level free list (PLAN amendment A5)
|
||||||
|
# ===========================================================================
|
||||||
|
#
|
||||||
|
# Same machine, same driver. Server at the M1 commit named per block,
|
||||||
|
# ReleaseFast. These are the numbers D7.4 said an M1 item owed.
|
||||||
|
|
||||||
|
[M1.1] churn gate — the doc-level free list
|
||||||
|
reproduce: node tests/e2e/churn.js --docs 40000 --doc-size 16k --index \
|
||||||
|
--mode delete-refill --rounds 6
|
||||||
|
node tests/e2e/churn.js --docs 40000 --doc-size 16k --index \
|
||||||
|
--mode update --multiple 5
|
||||||
|
baseline: the same harness against the end-of-Stage-2 binary (e416ad1),
|
||||||
|
which has no reclamation and no `multifora` section.
|
||||||
|
|
||||||
|
The harness is committed this time (`tests/e2e/churn.js`), which is half the
|
||||||
|
point of the block: D7.4's numbers were real and unrepeatable.
|
||||||
|
|
||||||
|
Stage 2 M1 target
|
||||||
|
delete half and refill, 6x 1.94x 1.94x <= 1.45x NOT MET
|
||||||
|
random $set over 5x the coll. 2.46x 2.46x <= 1.60x NOT MET
|
||||||
|
Both flat, drift +0.00x over the last three rounds.
|
||||||
|
|
||||||
|
D7.4 recorded 1.65x for the delete line. This harness reads 1.94x for the
|
||||||
|
*same binary* D7.4 was measured against the descendants of, so that gap is
|
||||||
|
the harness, not a regression: the ad-hoc version sampled ids to delete
|
||||||
|
blindly, which re-picks already-dead ids, deletes fewer than it inserts and
|
||||||
|
measures a collection that is quietly growing. The update line reproduces
|
||||||
|
D7.4 exactly (2.46 vs 2.47).
|
||||||
|
|
||||||
|
THE RATIO DID NOT MOVE AND THE MECHANISM WORKS. Both are true, and the
|
||||||
|
counters are what separate them:
|
||||||
|
|
||||||
|
round 6, update line: reclaimed 934.0MB dead 83.9MB
|
||||||
|
allocTail 1523.5MB freeReady 797.0MB
|
||||||
|
inUse 1.16x file/live 2.46x
|
||||||
|
|
||||||
|
Reclamation returned 934 MB over the run and the collection is occupying
|
||||||
|
1.16x its live data. What 2.46x measures is the data file's high-water mark,
|
||||||
|
and the file never shrinks. The mark is set once, in round 1, by the one
|
||||||
|
thing reclamation cannot avoid: a rebuild needs a whole second copy of the
|
||||||
|
live data before the first copy can be freed. 626 MB live + the garbage
|
||||||
|
standing at the moment it fires + 626 MB of copy is the number, and it is
|
||||||
|
reached before any free pool exists to build the copy out of.
|
||||||
|
|
||||||
|
So the floor for a rebuild-based design is ~2x, and no threshold reaches it.
|
||||||
|
Rebuilding earlier lowers the garbage term and raises nothing; rebuilding
|
||||||
|
later raises it. The plan anticipated this exact outcome and said what to do
|
||||||
|
about it, which is to write it down rather than tune: the remaining lever is
|
||||||
|
incremental compaction -- a doc-id-to-offset indirection layer, so a rebuild
|
||||||
|
moves documents without a second copy of everything. That is amendment A5's
|
||||||
|
successor and it is a milestone of its own, not a knob.
|
||||||
|
|
||||||
|
A second lever, cheaper and not attempted: give free space back to the
|
||||||
|
filesystem. `freeReady` stands at 797 MB with `allocTail` flat, so 52% of the
|
||||||
|
file is space the database owns and is not using. Returning the tail-adjacent
|
||||||
|
part of it needs the file never to shrink below what the fallback generation
|
||||||
|
references, which is a crash-safety argument and its own design pass.
|
||||||
|
|
||||||
|
What did change, and is the reason the mechanism is worth keeping:
|
||||||
|
|
||||||
|
delete-refill, 12k x 16 KiB reclaimed 1 MB -> 256 MB (six rounds)
|
||||||
|
|
||||||
|
The first figure is what reclamation achieved before `compact` was made to
|
||||||
|
checkpoint before it copies. A checkpoint is what reclaims and a checkpoint
|
||||||
|
is armed by log volume; a delete logs only an `_id`, so deleting half a
|
||||||
|
collection moved the log by a couple of megabytes, no checkpoint ran, and the
|
||||||
|
rebuild got there first every time and reset the window map it would have
|
||||||
|
used. The harness found that on its first serious run, which is the argument
|
||||||
|
for committing it.
|
||||||
|
|
||||||
|
[M1.2] churn gate — 200-byte documents
|
||||||
|
reproduce: node tests/e2e/churn.js --docs 150000 --doc-size 200 --index \
|
||||||
|
--mode delete-refill --rounds 4
|
||||||
|
Predicted in advance, in the plan, as a pass rather than a fault:
|
||||||
|
|
||||||
|
reclaimed 0.0 MB over four rounds, exactly as forecast.
|
||||||
|
ratio 3.93x on both the Stage 2 binary and M1 -- identical, flat.
|
||||||
|
|
||||||
|
Reclamation hands back whole system pages. A 16 KiB page on this machine
|
||||||
|
holds ~70 documents of 200 bytes and the chance that all 70 are dead at once
|
||||||
|
under uniform deletion is nil, so nothing is ever handed back. The forecast
|
||||||
|
said the ratio would not improve and the counters would show a mechanism
|
||||||
|
that correctly does nothing, rather than one that silently misfires; that is
|
||||||
|
what they show.
|
||||||
|
|
||||||
|
Read the ratio on this line with care. `live` counts document bytes, and at
|
||||||
|
200 bytes the two index trees are comparable in size to the documents
|
||||||
|
themselves -- the file is already 2.18x at load, before any churn. That
|
||||||
|
overhead is index structure, not slab garbage, and it is not what this gate
|
||||||
|
is about.
|
||||||
|
|
||||||
|
The payoff of window reclamation scales as doc_size / map_align, so a 4 KiB
|
||||||
|
system page (x86-64 Linux) reads four times better on the same code. Every
|
||||||
|
number in this file is Apple Silicon with 16 KiB pages.
|
||||||
|
|||||||
Reference in New Issue
Block a user