PLAN D2 makes the official specification suites the gate for command semantics; D7.6 asks for the harness to exist at M0 with a recorded baseline. This is that harness, pinned on both sides -- mongodb/specifications @ 615e0f9 and mongodb@7.5.0 -- because a scorecard is only comparable across milestones if a delta cannot be an upstream test change. It implements the unified format's Evaluating Matches algorithm as written, including the two rules that decide whether a pass is earned: extra keys are tolerated only in a root document, and numeric types compare flexibly. Anything unimplemented is a SKIP with a reason, never a pass, and the one assertion class not yet checked -- expectEvents, i.e. command monitoring -- is disclosed at the top of the scorecard so `pass` reads as an upper bound. First honest run: 131 pass, 161 fail, 195 skip over 175 files, zero timeouts. Getting there took four attempts, and the failures are documented in the README because each would have shipped a scorecard claiming a compatibility gap that did not exist. Two were genuine leaks in this runner (clients left open when a case timed out; clients registered for cleanup only after `await connect()`, plus abandoned cases still creating more). The third I misdiagnosed as machine load. The fourth attempt found the real cause: a leaked catalog lock in the engine, fixed separately, which alone accounts for the jump from 45 passes to 131. So the runner carries its own guards: per-operation CSOT timeouts so work is never abandoned, an active-handle census per file, an end-of-run tripwire for stray timers, a hard stop if the server dies rather than emitting hundreds of misleading ECONNREFUSED failures, and --skip/--limit for bisecting a run whose failures depend on position. The README states the rule plainly -- a long unbroken tail of timeouts is a harness bug until proven otherwise -- and the two commands that settle it. Also fixes bench-run.sh, which copied its report over bench-latest.txt unconditionally, including after a run that only warned -- so a degraded run could silently replace the baseline that PLAN D7.5 makes a milestone gate.
137 lines
6.7 KiB
Markdown
137 lines
6.7 KiB
Markdown
# MongoDB spec tests
|
|
|
|
PLAN D2 makes the official MongoDB JSON specification suites the gate for
|
|
command semantics: it turns "maximally compatible" into a concrete list of test
|
|
files rather than a judgement call. This directory holds the runner and the
|
|
committed scorecard.
|
|
|
|
```sh
|
|
bash tests/spec/fetch.sh # pinned suites (~175 files, gitignored)
|
|
zig build # the runner spawns this binary
|
|
node tests/spec/run.js # run everything
|
|
node tests/spec/run.js --scorecard # ... and rewrite scorecard.txt
|
|
node tests/spec/run.js --file find.json --verbose
|
|
node tests/spec/run.js --url mongodb://127.0.0.1:27020 # use a server you started
|
|
```
|
|
|
|
## What is pinned, and why both halves matter
|
|
|
|
- **Suites**: `mongodb/specifications` @ `615e0f9`, in `fetch.sh`.
|
|
- **Driver**: `mongodb@7.5.0`, via `tests/e2e/package-lock.json`.
|
|
|
|
A scorecard is only comparable across milestones if both are pinned — otherwise
|
|
a delta could be an upstream test change rather than an engine change. Bump
|
|
either one in its own commit and re-record the scorecard in that same commit.
|
|
|
|
The suites are fetched rather than vendored: they are someone else's corpus,
|
|
upstream rewrites them wholesale, and a pinned commit gives the same
|
|
reproducibility without putting them in this repo's history.
|
|
|
|
## Scope
|
|
|
|
`source/crud/tests/unified/` — 175 files. The aggregate tests live there too
|
|
(`aggregate*.json`), so this one directory is PLAN M0's "crud + aggregate".
|
|
|
|
The runner implements the unified test format's **Evaluating Matches**
|
|
algorithm as written in the spec, including the two rules that decide whether a
|
|
result is a real pass:
|
|
|
|
- extra keys in the actual document are tolerated **only** in a root document;
|
|
- numeric types (int32 / int64 / double) compare flexibly.
|
|
|
|
Supported: `client`/`database`/`collection` entities, `initialData`,
|
|
`outcome`, `expectError` (code, codeName, contains, labels, errorResponse),
|
|
`saveResultAsEntity`, `runOnRequirements` gating, and the `$$type`,
|
|
`$$exists`, `$$unsetOrMatches`, `$$matchesEntity`, `$$matchesHexBytes`
|
|
operators.
|
|
|
|
**Not asserted yet: `expectEvents`** (command monitoring). Those assertions are
|
|
about the command shape the driver emits rather than result semantics. Ignoring
|
|
them lets some cases pass that a complete runner would fail, so **treat `pass`
|
|
as an upper bound** until M1 wires events up. This is stated again at the top of
|
|
`scorecard.txt` so the number is never read out of context.
|
|
|
|
Not supported, each reported as SKIP with a reason and never as PASS: session
|
|
and bucket entities (M4 / GridFS), `failPoint`, client-side encryption,
|
|
`testRunner` operations, and any operation or matcher the runner does not know.
|
|
|
|
## Reading the scorecard
|
|
|
|
`scorecard.txt` records the totals, a per-file breakdown, and every
|
|
non-passing case with its reason. The distinction that matters:
|
|
|
|
- **FAIL** — the engine answered, and answered differently from the spec. Real
|
|
work. An operation that never answered inside `--op-timeout-ms` (default 3 s,
|
|
enforced by the driver itself via CSOT `timeoutMS`) is also a FAIL, because
|
|
"no answer" is a result. There is a second, much longer `--case-timeout-ms`
|
|
backstop for a hang the driver cannot see; if it ever fires, treat the run
|
|
with suspicion — see the trap below.
|
|
- **SKIP** — nobody claims anything. Either the suite needs a feature whose
|
|
milestone has not landed, or the runner does not implement it yet.
|
|
|
|
M0's gate (PLAN D7.6) is only that the harness exists and the baseline is
|
|
recorded. **A red baseline is the expected state**, so `run.js` exits 0 as long
|
|
as it ran; it is a measuring tool, not a pass/fail gate. Later milestones move
|
|
the numbers, and each one commits the new scorecard (PLAN D9).
|
|
|
|
## A trap worth knowing about: the harness can invent failures
|
|
|
|
The first baseline attempt reported ~77 timeout FAILs that did not exist. Every
|
|
case from one point onward timed out, while a `ping` from a separate process
|
|
answered instantly — which read convincingly as a server-side wedge, and was
|
|
not.
|
|
|
|
The cause was in this runner. `buildEntities` opened `MongoClient`s, and a case
|
|
that timed out before it returned left them unclosed; each one keeps a
|
|
connection pool and a heartbeat timer. Once enough accumulated, Node's event
|
|
loop was starved badly enough that the per-case timer fired before operations
|
|
could finish. Then every later case "failed".
|
|
|
|
Two things guard it now: per-test clients are owned by the caller and closed
|
|
unconditionally, including on a partial failure; and the run ends by checking
|
|
how many timers are still active, warning loudly if the answer is more than a
|
|
handful.
|
|
|
|
The general rule, since it will come up again: **a run with a long unbroken tail
|
|
of timeouts is a harness bug until proven otherwise.** Confirm it by running the
|
|
first timing-out file on its own — if it passes in isolation, the failures are
|
|
this runner's, not the engine's.
|
|
|
|
### ... but the third time it was the engine
|
|
|
|
A later attempt produced 166 timeout FAILs starting at file 70. I first blamed
|
|
machine load — a concurrent `zig build test` against a then-3-second budget —
|
|
and that was **wrong**. The evidence against it: the collapse reproduced on an
|
|
idle machine, at the same file, with a 10 s budget.
|
|
|
|
The actual cause was a leaked catalog lock in the engine, and it is worth
|
|
knowing how it hid. `db-aggregate.json` sends `{aggregate: 1}`, which names no
|
|
collection; dispatch resolved the namespace after taking the catalog lock and
|
|
bailed with a plain `return`, holding it shared forever. A leaked *shared* lock
|
|
is invisible to readers, so the server stayed perfectly responsive — an external
|
|
prober got `ok 15ms` right through the hang — and only the next write that had
|
|
to take the catalog exclusive to create a collection blocked. The failure
|
|
therefore surfaced one file later, on a different connection, as a client-side
|
|
timeout with nothing pointing at its cause.
|
|
|
|
Two lessons for using this runner:
|
|
|
|
- **A healthy-looking server does not exonerate the engine.** Probe with the
|
|
operation that is actually stuck, not with `ping`.
|
|
- **The driver's own command log is the fastest way in.** It showed an insert
|
|
sitting for exactly `socketTimeoutMS` against an idle engine, which is what
|
|
turned a week-long-looking mystery into a five-line fix:
|
|
|
|
```sh
|
|
MONGODB_LOG_COMMAND=debug MONGODB_LOG_PATH=stderr \
|
|
node tests/spec/run.js --skip 68 --limit 2 2>drv.log
|
|
```
|
|
|
|
`--skip`/`--limit` exist for exactly this: the collapse reduced to a
|
|
reproducible two-file window, which is what made it tractable.
|
|
|
|
Still worth recording the baseline on an otherwise idle machine, and do not
|
|
tighten `--op-timeout-ms` to make a run finish sooner — a tight budget turns
|
|
load into apparent engine failures, which is how I misdiagnosed this once
|
|
already.
|