tests/spec: MongoDB spec-test runner and the M0 scorecard

PLAN D2 makes the official specification suites the gate for command semantics;
D7.6 asks for the harness to exist at M0 with a recorded baseline. This is that
harness, pinned on both sides -- mongodb/specifications @ 615e0f9 and
mongodb@7.5.0 -- because a scorecard is only comparable across milestones if a
delta cannot be an upstream test change.

It implements the unified format's Evaluating Matches algorithm as written,
including the two rules that decide whether a pass is earned: extra keys are
tolerated only in a root document, and numeric types compare flexibly. Anything
unimplemented is a SKIP with a reason, never a pass, and the one assertion class
not yet checked -- expectEvents, i.e. command monitoring -- is disclosed at the
top of the scorecard so `pass` reads as an upper bound.

First honest run: 131 pass, 161 fail, 195 skip over 175 files, zero timeouts.

Getting there took four attempts, and the failures are documented in the README
because each would have shipped a scorecard claiming a compatibility gap that
did not exist. Two were genuine leaks in this runner (clients left open when a
case timed out; clients registered for cleanup only after `await connect()`,
plus abandoned cases still creating more). The third I misdiagnosed as machine
load. The fourth attempt found the real cause: a leaked catalog lock in the
engine, fixed separately, which alone accounts for the jump from 45 passes to
131.

So the runner carries its own guards: per-operation CSOT timeouts so work is
never abandoned, an active-handle census per file, an end-of-run tripwire for
stray timers, a hard stop if the server dies rather than emitting hundreds of
misleading ECONNREFUSED failures, and --skip/--limit for bisecting a run whose
failures depend on position. The README states the rule plainly -- a long
unbroken tail of timeouts is a harness bug until proven otherwise -- and the two
commands that settle it.

Also fixes bench-run.sh, which copied its report over bench-latest.txt
unconditionally, including after a run that only warned -- so a degraded run
could silently replace the baseline that PLAN D7.5 makes a milestone gate.
This commit is contained in:
2026-08-03 18:58:01 +03:00
parent e2c25a986b
commit 90de7820da
7 changed files with 1487 additions and 6 deletions

136
tests/spec/README.md Normal file
View File

@@ -0,0 +1,136 @@
# MongoDB spec tests
PLAN D2 makes the official MongoDB JSON specification suites the gate for
command semantics: it turns "maximally compatible" into a concrete list of test
files rather than a judgement call. This directory holds the runner and the
committed scorecard.
```sh
bash tests/spec/fetch.sh # pinned suites (~175 files, gitignored)
zig build # the runner spawns this binary
node tests/spec/run.js # run everything
node tests/spec/run.js --scorecard # ... and rewrite scorecard.txt
node tests/spec/run.js --file find.json --verbose
node tests/spec/run.js --url mongodb://127.0.0.1:27020 # use a server you started
```
## What is pinned, and why both halves matter
- **Suites**: `mongodb/specifications` @ `615e0f9`, in `fetch.sh`.
- **Driver**: `mongodb@7.5.0`, via `tests/e2e/package-lock.json`.
A scorecard is only comparable across milestones if both are pinned — otherwise
a delta could be an upstream test change rather than an engine change. Bump
either one in its own commit and re-record the scorecard in that same commit.
The suites are fetched rather than vendored: they are someone else's corpus,
upstream rewrites them wholesale, and a pinned commit gives the same
reproducibility without putting them in this repo's history.
## Scope
`source/crud/tests/unified/` — 175 files. The aggregate tests live there too
(`aggregate*.json`), so this one directory is PLAN M0's "crud + aggregate".
The runner implements the unified test format's **Evaluating Matches**
algorithm as written in the spec, including the two rules that decide whether a
result is a real pass:
- extra keys in the actual document are tolerated **only** in a root document;
- numeric types (int32 / int64 / double) compare flexibly.
Supported: `client`/`database`/`collection` entities, `initialData`,
`outcome`, `expectError` (code, codeName, contains, labels, errorResponse),
`saveResultAsEntity`, `runOnRequirements` gating, and the `$$type`,
`$$exists`, `$$unsetOrMatches`, `$$matchesEntity`, `$$matchesHexBytes`
operators.
**Not asserted yet: `expectEvents`** (command monitoring). Those assertions are
about the command shape the driver emits rather than result semantics. Ignoring
them lets some cases pass that a complete runner would fail, so **treat `pass`
as an upper bound** until M1 wires events up. This is stated again at the top of
`scorecard.txt` so the number is never read out of context.
Not supported, each reported as SKIP with a reason and never as PASS: session
and bucket entities (M4 / GridFS), `failPoint`, client-side encryption,
`testRunner` operations, and any operation or matcher the runner does not know.
## Reading the scorecard
`scorecard.txt` records the totals, a per-file breakdown, and every
non-passing case with its reason. The distinction that matters:
- **FAIL** — the engine answered, and answered differently from the spec. Real
work. An operation that never answered inside `--op-timeout-ms` (default 3 s,
enforced by the driver itself via CSOT `timeoutMS`) is also a FAIL, because
"no answer" is a result. There is a second, much longer `--case-timeout-ms`
backstop for a hang the driver cannot see; if it ever fires, treat the run
with suspicion — see the trap below.
- **SKIP** — nobody claims anything. Either the suite needs a feature whose
milestone has not landed, or the runner does not implement it yet.
M0's gate (PLAN D7.6) is only that the harness exists and the baseline is
recorded. **A red baseline is the expected state**, so `run.js` exits 0 as long
as it ran; it is a measuring tool, not a pass/fail gate. Later milestones move
the numbers, and each one commits the new scorecard (PLAN D9).
## A trap worth knowing about: the harness can invent failures
The first baseline attempt reported ~77 timeout FAILs that did not exist. Every
case from one point onward timed out, while a `ping` from a separate process
answered instantly — which read convincingly as a server-side wedge, and was
not.
The cause was in this runner. `buildEntities` opened `MongoClient`s, and a case
that timed out before it returned left them unclosed; each one keeps a
connection pool and a heartbeat timer. Once enough accumulated, Node's event
loop was starved badly enough that the per-case timer fired before operations
could finish. Then every later case "failed".
Two things guard it now: per-test clients are owned by the caller and closed
unconditionally, including on a partial failure; and the run ends by checking
how many timers are still active, warning loudly if the answer is more than a
handful.
The general rule, since it will come up again: **a run with a long unbroken tail
of timeouts is a harness bug until proven otherwise.** Confirm it by running the
first timing-out file on its own — if it passes in isolation, the failures are
this runner's, not the engine's.
### ... but the third time it was the engine
A later attempt produced 166 timeout FAILs starting at file 70. I first blamed
machine load — a concurrent `zig build test` against a then-3-second budget —
and that was **wrong**. The evidence against it: the collapse reproduced on an
idle machine, at the same file, with a 10 s budget.
The actual cause was a leaked catalog lock in the engine, and it is worth
knowing how it hid. `db-aggregate.json` sends `{aggregate: 1}`, which names no
collection; dispatch resolved the namespace after taking the catalog lock and
bailed with a plain `return`, holding it shared forever. A leaked *shared* lock
is invisible to readers, so the server stayed perfectly responsive — an external
prober got `ok 15ms` right through the hang — and only the next write that had
to take the catalog exclusive to create a collection blocked. The failure
therefore surfaced one file later, on a different connection, as a client-side
timeout with nothing pointing at its cause.
Two lessons for using this runner:
- **A healthy-looking server does not exonerate the engine.** Probe with the
operation that is actually stuck, not with `ping`.
- **The driver's own command log is the fastest way in.** It showed an insert
sitting for exactly `socketTimeoutMS` against an idle engine, which is what
turned a week-long-looking mystery into a five-line fix:
```sh
MONGODB_LOG_COMMAND=debug MONGODB_LOG_PATH=stderr \
node tests/spec/run.js --skip 68 --limit 2 2>drv.log
```
`--skip`/`--limit` exist for exactly this: the collapse reduced to a
reproducible two-file window, which is what made it tractable.
Still worth recording the baseline on an otherwise idle machine, and do not
tighten `--op-timeout-ms` to make a run finish sooner — a tight budget turns
load into apparent engine failures, which is how I misdiagnosed this once
already.