0034da4293ae344fcfb96dc0a22e049a11c10636
8 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e65da2740e |
plan: what the event assertions found
Four defects and one deliberate non-fix, recorded where the milestone can see them rather than only in five commit messages. The one worth carrying forward: the runner dropped `collectionOptions`, so every "unacknowledged write" case in the corpus had been running an acknowledged write and passing, because the two produce results the expectation accepts either way. Only the wire distinguished them and nothing read the wire. Also states plainly that scorecards from before this are not comparable with ones after, and why `bypassDocumentValidation` is left failing: both available refusals are worse than four undeserved entries in the fail column, and one of them is the shape that turns a scorecard into flattery. |
||
|
|
e66030f43e |
plan: the eight bugs cleared before the free list
Four of them were not in the plan that started this work -- they were found by tests written for the three that were, which is M0's gate lesson arriving a milestone early. The record has to match what happened, or the next session reads a milestone that looks like it went as designed. Also records what nobody is fixing yet: `rebuild_collection` frees pages under only the collection's lock while a concurrent checkpoint may have snapshotted a catalog that claims them, and the `seq` retry cannot see it because a rebuild appends no log record. Written down so the free list does not add a second instance of the same shape. |
||
| f2844e7894 |
cursors: server-side cursors for find, aggregate and the listing commands
Every reply came back in a single batch with `cursor.id = 0`, `getMore` was a
stub answering an empty `nextBatch` on the literal namespace `test.$cmd`, and
nothing read `batchSize`. That caps the useful collection size at what fits in
one 48 MiB message, which is the opposite of the tens-of-GB target and the
reason M0 made whole-index scans stream: the streaming candidate generator
existed with no consumer that could suspend.
## What a cursor is allowed to remember
A cursor holds no lock between requests, so everything it saves has to survive
arbitrary concurrent mutation. Nothing here is a pointer, and the two things
that look like stable addresses are not: `reset_tree` re-creates node ids 0 and
1 as different nodes, and `rebuild_collection` moves every document. Three
sources, chosen by query shape, each with a different memory contract:
- **stream** -- an index-ordered walk resumed from a `(key, off)` anchor plus a
`(leaf, slot)` hint. O(key) state, so this is what lets a cursor walk a
collection larger than memory. Survives a rebuild, because a repack changes
no key.
- **offsets** -- the matched slab offsets a narrowed plan already materialized,
8 bytes each. Killed by a rebuild with `QueryPlanKilled`, because those
offsets now name unrelated bytes.
- **buffered** -- canonical BSON copies, for a sort no index provides and for
aggregate/listing output. Depends on nothing, which is what lets a listing
hold a cursor over a `$cmd.*` namespace no collection backs.
`Collection.layout_epoch` and `Index.epoch` are the invalidation tokens, both
checked as error returns rather than assertions since a client reaches them by
keeping a cursor open across maintenance.
## Resume
`resume_forward`/`resume_reverse` are O(1) while the hint holds and fall back to
an exact-order band walk bounded by `resume_walk_max`. Without the hint, `seek`
lands at the *start* of an equal-key band, so `sort({status: 1})` over three
distinct values across 10M documents would cost ~5e10 comparisons to drain.
Two hazards found by draining a collection while writing to it, neither
predictable from reading the code:
- A deleted anchor must resume at its *band position*, or the rest of an
equal-key band is silently dropped -- most of the collection on a
low-cardinality index. Hence `band_index`.
- On a **unique** index a same-key entry can only be the anchor rewritten, so
resuming at it returned updated documents twice. Observed as duplicate `_id`s
while updating underneath a drain.
## Protocol
Measured against mongod 8.3.7 rather than recalled, which corrected three
assumptions: a bare `getMore` does *not* inherit the find's `batchSize` (4998 of
5000 documents come back), a namespace mismatch is `Unauthorized` (13) not
`CursorNotFound`, and `CursorInUse` is 143 not 12051.
`internalQueryFindCommandBatchSize` is 101, `cursorTimeoutMillis` 600000,
`clientCursorMonitorFrequencySecs` 4.
The rule everything follows is **never look ahead**: a batch that met its target
leaves the cursor open even when the source is in fact exhausted, so four
documents at `batchSize: 2` take three commands. `limit` acts as an EOF source,
which is what makes `batchSize == limit` close in one round trip. `skip` is
consumed once. `batchSize: 0` returns an empty batch with a live cursor.
Cursor ids are `(nonce << 20) | slot`, always positive. The nonce is not
decoration: without it a recycled slot serves one client another's documents.
Cursors are not connection-pinned, since the driver spec allows a `getMore` on
any connection to the same server; they end at exhaustion, `killCursors`, or the
idle sweep (a second monitor fiber, separate from the TTL one because the
cadences differ by an order of magnitude and a TTL failure must not stop
reclamation). The registry is fixed-capacity and evicts the least recently used
cursor, whose client sees the same 43 an idle timeout gives.
Fixed alongside, because cursors are what expose them:
- `listCollections` reported `"<db>."` with an *empty* collection part, which
makes the driver throw client-side -- so it would have broken the moment its
cursor stopped being id 0. Now `<db>.$cmd.listCollections`, as mongod uses.
- `count` ignored `skip` and `limit` entirely.
- `wire.end_message` now bounds a reply by the 48 MiB we advertise rather than
by `maxInt(u32)`; a reply past what we told the client to expect is not a large
reply, it is a desynchronized connection.
- Two `codeName` strings were wrong: 72 is `InvalidOptions` (MongoDB has no
`InvalidArgument`), and 40324 reports as `Location40324`.
## Verification
Unit 160/160 in ReleaseFast and ReleaseSafe; `tests/e2e/e2e7.js` adds 86 cursor
checks across five phases (batching/lifecycle/errors, streaming across churn,
aggregate+listings+count, expiry+capacity, restart) and is self-contained
because cursor behaviour is only observable with non-default flags. No
regressions: e2e 49, e2e3 16, e2e4 17, e2e2 2, e2e6 72. Spec 168 pass / 124
fail, +5 against the previous scorecard.
Mutation-checked, per the repo's second ground rule: `hint_slot + 1`, the
`band_index` off-by-one, both epoch bumps, the id nonce, the at-least-one-
document rule, and `stream_shape` returning null each turn the intended test
red. One claim was withdrawn rather than kept -- swapping `std.mem.order` for
`cmp_prefix` in the band walk changes nothing observable, so the comment now
says so instead of asserting a check that does not hold.
|
|||
| d597597a4c |
plan/results: record the two post-gate CRUD corrections
The gate results file said the replacement-style-update gap was found and not fixed, and PLAN said it was left alone. Both were true when written and are not now, so a reader would take the M0 scorecard for the current one. The M0 figures stay as measured -- they are the gate result -- with a pointer to tests/spec/scorecard.txt, which always holds the current number. |
|||
| 504179acd1 |
results: the M0 gates, measured
PLAN D7's six items, with the numbers and the command that reproduces each in
tests/e2e/results/m0-gates.txt. Unit tests green in both optimize modes, the
whole e2e matrix green, the spec scorecard byte-identical at 131/161/195, and
the large smoke run at the scale D7.3 asked for:
21.47 GB collection (1,310,720 x 16 KiB)
data file 21.75 GB (+1.3% over the documents)
log after the load 2.5 MB (checkpoints reclaim it)
kill -9 then reopen 0.5 s (0.5 s at 4 GB too -- flat)
RSS after reopen 237 MB (1.1% of the data)
count after restart 1,310,720 last document byte-intact
acked writes after kill 200/200
That is the milestone's claim, measured: an open costs the working set rather
than the size of the database. Before M0 the same measurement was 523 MB
resident for a 512 MB database, because recovering each document's `_id` meant
reading every document at open.
Two gates need reading rather than a tick, and m0-gates.txt says so where a
reader would otherwise take a tick for granted.
The churn gate settles at 1.65x live data (delete-heavy) to 2.47x
(update-heavy), flat, above the ~1.3x amendment A2 hoped for. Rebuild-only
reclamation cannot reach that: it needs a whole second copy of the live data
before the first can be freed. The gate existed to decide whether doc-level
free lists are needed after M0, and that is the answer.
Benchmark parity holds for every read and latency row inside the run-to-run
spread, and bulk insert regresses 24% (732 -> 555 MB/s), reproducibly across
three runs. Risk 1 as written: document bytes now reach the disk uncompressed
on top of the LZ4 log. createIndex improves 62% from the same change.
Three measurement bugs fixed while running the gates, because each would have
put a false number in the README:
- `compare-run.sh` measured "db on disk" as `du` of the log alone against
`du` of mongod's whole dbpath. It reported 20 MB for a 1 GB collection --
the documents had moved to <db>.data. Honest figure, measured: 914 MB of
allocated blocks against mongod's compressed 85 MB.
- `big.js` counted "compaction events" as "the log shrank", which is a
*checkpoint* now. It claimed 12 compaction rewrites during a pure insert
load, which has no garbage to compact.
- `big.js` labelled peak RSS "in-memory engine: docs live in RAM" and its
summary said the collection was held "fully in RAM". Both were true of the
engine this milestone replaced.
README: the storage section described an all-in-RAM engine; the comparison
table mixed one old run's body with three new rows; and `findOne({_id})` was
documented as a full scan for integer ids, which the ordered `_id_` index made
false (2 ms against 55 s for a scan of the same 21.5 GB collection). The table
is now best-of-three for both servers, with the measured variance stated, since
two runs of the same binary moved the sub-10 ms rows by 27-51%.
|
|||
| 90de7820da |
tests/spec: MongoDB spec-test runner and the M0 scorecard
PLAN D2 makes the official specification suites the gate for command semantics; D7.6 asks for the harness to exist at M0 with a recorded baseline. This is that harness, pinned on both sides -- mongodb/specifications @ 615e0f9 and mongodb@7.5.0 -- because a scorecard is only comparable across milestones if a delta cannot be an upstream test change. It implements the unified format's Evaluating Matches algorithm as written, including the two rules that decide whether a pass is earned: extra keys are tolerated only in a root document, and numeric types compare flexibly. Anything unimplemented is a SKIP with a reason, never a pass, and the one assertion class not yet checked -- expectEvents, i.e. command monitoring -- is disclosed at the top of the scorecard so `pass` reads as an upper bound. First honest run: 131 pass, 161 fail, 195 skip over 175 files, zero timeouts. Getting there took four attempts, and the failures are documented in the README because each would have shipped a scorecard claiming a compatibility gap that did not exist. Two were genuine leaks in this runner (clients left open when a case timed out; clients registered for cleanup only after `await connect()`, plus abandoned cases still creating more). The third I misdiagnosed as machine load. The fourth attempt found the real cause: a leaked catalog lock in the engine, fixed separately, which alone accounts for the jump from 45 passes to 131. So the runner carries its own guards: per-operation CSOT timeouts so work is never abandoned, an active-handle census per file, an end-of-run tripwire for stray timers, a hard stop if the server dies rather than emitting hundreds of misleading ECONNREFUSED failures, and --skip/--limit for bisecting a run whose failures depend on position. The README states the rule plainly -- a long unbroken tail of timeouts is a harness bug until proven otherwise -- and the two commands that settle it. Also fixes bench-run.sh, which copied its report over bench-latest.txt unconditionally, including after a run that only warned -- so a degraded run could silently replace the baseline that PLAN D7.5 makes a milestone gate. |
|||
| aee23cb028 |
index/db: enforce _id uniqueness through the _id_ index
_id uniqueness was a `coll.docs.contains` probe. The docs hashmap is going
away (PLAN A3), so it has to move to the _id_ tree -- and the tree answers
better, because it is keyed on bson.encode_key, which is canonical where
serialize_value is not. int32 1, int64 1 and double 1.0 are now one _id, as
they are in MongoDB (A4).
_id_ is built and checked first, so a write violating both it and a unique
secondary reports _id_, which is what MongoDB reports. It returns
error.DuplicateKey with `dup_index` left null, which is exactly what
commands.zig's E11000 rendering already treats as "the _id_ index", so the
wire-visible message is unchanged and that file needed no edit.
check_unique's exclude-self became optional and is null on an insert. That was
a latent bug of its own: a replace must ignore its own existing entries, but an
insert has none, and passing the document's id there hides a collision whose
entry carries that same id -- precisely the case _id_ exists to catch. Only
_id_ could reach it, since a secondary collision is between different
documents.
Two corrections found while doing this, both worth reading:
PLAN A4 claimed a database already holding {_id: int32 1} and {_id: int64 1}
loses one on reopen. It does not. Replay evicts through the docs map, keyed on
serialize_value, so both survive; the tree is bulk-built afterwards with
enforcement off, which tolerates duplicate keys and warns. The loss arrives
only with the commit that drops the map, and that is where it needs a
pre-flight scan. Amended.
dispatch_insert asserted only `ok: 1`, but a rejected document comes back as a
writeError alongside it -- so the mixed-type corpus silently shrank from ten
documents to nine when _id_ became unique, and every test over it still passed.
The helper now rejects writeErrors and asserts the inserted count; it caught
the shrink immediately. The corpus keeps an int64 _id on a distinct value, and
the collision it used to stand in for is asserted directly.
Also adds Index.lookup_exact, which the commands that currently probe the docs
map will need. Exact byte equality rather than cmp_prefix, because {a: 1}'s
encoding is a proper prefix of {a: 1, b: 2}'s and a prefix match would claim a
document is present when it is not.
Mutation-checked, all three red: unique=false on id_index; exclude=id_key on
insert; eql -> cmp_prefix in lookup_exact.
|
|||
| 13d7b79f2c |
plan: amend the M0 decision record for copy-on-write; add AGENTS.md
The M0 implementation review found three of D1-D9 wrong or incomplete. The originals stay in place with pointers to a new amendments section, so a later session can see what changed rather than reading a rewritten history. A1: D4 as written is unsound. The doc and overflow slabs are append-only, so replay repairs them, but B+tree node pages are mutated in place -- after a crash the file holds an arbitrary mix of written-back and not-written-back pages, and once D6.3 truncates the log the data file is the only copy below the watermark. A half-persisted tree is unrecoverable. So the checkpoint needs shadow paging: no page below the last watermark's allocation mark is ever stored into, and the watermark write is the atomic switch. Knock-on: node ids cannot be page numbers, because Node.parent/next/prev are back-pointers by id and copy-on-write would cascade; an in-RAM id->page table per index keeps every persisted id at its current width and gives COW one pointer to fix. A2: follows from A1 -- a page free list is a prerequisite, not the defense-in-depth D6.2 assumed, because COW abandons every page it touches in every epoch. The churn gate stays, retargeted at document garbage. A3: section 5 step 4's trap was misidentified. store_record already copies keys, so "entries must own their key bytes" is work that does not need doing. The real problem is that a leaf record has nowhere to put a slab offset; the resolution is to make the payload that offset, in every index, and delete Entry.id rather than re-own it. A4: making _id_ a unique index keys uniqueness on the canonical encode_key rather than serialize_value, so int32 1 / int64 1 / double 1.0 collide as they do in MongoDB -- a compatibility improvement, with a documented one-way migration hazard for a database that already holds two such documents. Also records that Engine.seq is never restored on open (harmless today, silent data loss once a watermark exists), and two bugs the new spec harness found. AGENTS.md carries the same rules into the operating guide: ground rules grow from 7 to 9, and the old rule 6 is corrected with a note saying why. |