Files
MultiforaDB/tests/spec
A.Shakhmatov 97e3e3a556 tests/spec: assert expectEvents
The headline is not the delta, it is that `pass` changed meaning. 354 of the
487 cases declare `expectEvents` and until now the runner read none of them,
so a case could send the wrong command entirely and still be counted a pass
as long as the *result* came back right. The old column was an upper bound by
construction. 193/99/195 becomes 159/133/195, and the two numbers are not
comparable.

Two rules decide how far the assertion reaches, both taken from the spec
rather than from what would be convenient:

  - `command` and `reply` match as *root* documents
    (unified-test-format.md:1020-1022, :1037-1039). The driver hangs `lsid`,
    `$db` and `maxTimeMS` off nearly everything it sends; as nested documents
    essentially the whole corpus would fail on keys no expectation was ever
    written to mention, and the number would say nothing.
  - the event list is exact in number and order, not a prefix
    (unified-test-format.md:3088-3091). 23 cases expect an empty list and a
    prefix rule would pass every one of them without looking.

The assertion runs after the operations, so a wrong result is still reported
as a wrong result rather than being masked, and after the listeners are
disabled, so the teardown's own commands cannot reach the buffer.

`cmap` and `sdam` event types, `ignoreExtraEvents`, and any event field beyond
`command`/`reply`/`commandName`/`databaseName` are reported unsupported at the
point of assertion. None occurs in this corpus -- all 354 blocks are
`eventType: command`, carrying 349 `commandStartedEvent` and 6
`commandSucceededEvent` -- so nothing is being quietly waived.

All 34 newly-failing cases, triaged. Not one is a wrong answer from the
engine; every one is a command the driver never sent:

  - 22x `command.writeConcern: missing` -- runner gap, and the sharpest thing
    this commit found. `buildEntities` drops `collectionOptions` on the floor,
    so `writeConcern: {w: 0}` never reached the driver and every "unacknowledged
    write" case in the corpus has been running an acknowledged write. They
    passed because the results of the two agree. This is precisely the class of
    error the instrument was built to find, and it was invisible to the result
    column.
  - 5x `command.sort.<key>: missing` -- runner gap. The driver holds a sort as
    a JS `Map` (lib/sort.js), so `Object.keys` on it is empty and the matcher
    reports every expected key as absent. Measured, not guessed: EJSON prints a
    `Map` exactly like a document, which is why the dump looks correct.
  - 4x `command.bypassDocumentValidation: missing` -- unclassified. The option
    is absent from the wire for the `false` cases; the driver only forwards it
    when true on some paths (lib/operations/find_and_modify.js:19), and whether
    the runner also drops it has not been established.
  - 2x `command.comment: missing` on getMore -- server gap, most likely. The
    driver gates it on `maxWireVersion >= 9` (lib/operations/get_more.js:43)
    and this engine advertises 8 while reporting itself as 4.4.0, which is
    wire 9. The inconsistency is ours.
  - 1x `command.maxTimeMS: expected 6000, got 10000` -- the CSOT rewrite, dealt
    with in the next commit.

Each of those gets its own commit, and none of them is fixed here: a check and
the fix for what the check caught do not belong in one change.
2026-08-09 13:11:43 +03:00
..
2026-08-09 13:11:43 +03:00
2026-08-09 13:11:43 +03:00

MongoDB spec tests

PLAN D2 makes the official MongoDB JSON specification suites the gate for command semantics: it turns "maximally compatible" into a concrete list of test files rather than a judgement call. This directory holds the runner and the committed scorecard.

bash tests/spec/fetch.sh                  # pinned suites (~175 files, gitignored)
zig build                                 # the runner spawns this binary
node tests/spec/run.js                    # run everything
node tests/spec/run.js --scorecard        # ... and rewrite scorecard.txt
node tests/spec/run.js --file find.json --verbose
node tests/spec/run.js --url mongodb://127.0.0.1:27020   # use a server you started

What is pinned, and why both halves matter

  • Suites: mongodb/specifications @ 615e0f9, in fetch.sh.
  • Driver: mongodb@7.5.0, via tests/e2e/package-lock.json.

A scorecard is only comparable across milestones if both are pinned — otherwise a delta could be an upstream test change rather than an engine change. Bump either one in its own commit and re-record the scorecard in that same commit.

The suites are fetched rather than vendored: they are someone else's corpus, upstream rewrites them wholesale, and a pinned commit gives the same reproducibility without putting them in this repo's history.

Scope

source/crud/tests/unified/ — 175 files. The aggregate tests live there too (aggregate*.json), so this one directory is PLAN M0's "crud + aggregate".

The runner implements the unified test format's Evaluating Matches algorithm as written in the spec, including the two rules that decide whether a result is a real pass:

  • extra keys in the actual document are tolerated only in a root document;
  • numeric types (int32 / int64 / double) compare flexibly.

Supported: client/database/collection entities, initialData, outcome, expectError (code, codeName, contains, labels, errorResponse), saveResultAsEntity, runOnRequirements gating, and the $$type, $$exists, $$unsetOrMatches, $$matchesEntity, $$matchesHexBytes operators.

Not asserted yet: expectEvents (command monitoring). Those assertions are about the command shape the driver emits rather than result semantics. Ignoring them lets some cases pass that a complete runner would fail, so treat pass as an upper bound until M1 wires events up. This is stated again at the top of scorecard.txt so the number is never read out of context.

Not supported, each reported as SKIP with a reason and never as PASS: session and bucket entities (M4 / GridFS), failPoint, client-side encryption, testRunner operations, and any operation or matcher the runner does not know.

Reading the scorecard

scorecard.txt records the totals, a per-file breakdown, and every non-passing case with its reason. The distinction that matters:

  • FAIL — the engine answered, and answered differently from the spec. Real work. An operation that never answered inside --op-timeout-ms (default 3 s, enforced by the driver itself via CSOT timeoutMS) is also a FAIL, because "no answer" is a result. There is a second, much longer --case-timeout-ms backstop for a hang the driver cannot see; if it ever fires, treat the run with suspicion — see the trap below.
  • SKIP — nobody claims anything. Either the suite needs a feature whose milestone has not landed, or the runner does not implement it yet.

M0's gate (PLAN D7.6) is only that the harness exists and the baseline is recorded. A red baseline is the expected state, so run.js exits 0 as long as it ran; it is a measuring tool, not a pass/fail gate. Later milestones move the numbers, and each one commits the new scorecard (PLAN D9).

A trap worth knowing about: the harness can invent failures

The first baseline attempt reported ~77 timeout FAILs that did not exist. Every case from one point onward timed out, while a ping from a separate process answered instantly — which read convincingly as a server-side wedge, and was not.

The cause was in this runner. buildEntities opened MongoClients, and a case that timed out before it returned left them unclosed; each one keeps a connection pool and a heartbeat timer. Once enough accumulated, Node's event loop was starved badly enough that the per-case timer fired before operations could finish. Then every later case "failed".

Two things guard it now: per-test clients are owned by the caller and closed unconditionally, including on a partial failure; and the run ends by checking how many timers are still active, warning loudly if the answer is more than a handful.

The general rule, since it will come up again: a run with a long unbroken tail of timeouts is a harness bug until proven otherwise. Confirm it by running the first timing-out file on its own — if it passes in isolation, the failures are this runner's, not the engine's.

... but the third time it was the engine

A later attempt produced 166 timeout FAILs starting at file 70. I first blamed machine load — a concurrent zig build test against a then-3-second budget — and that was wrong. The evidence against it: the collapse reproduced on an idle machine, at the same file, with a 10 s budget.

The actual cause was a leaked catalog lock in the engine, and it is worth knowing how it hid. db-aggregate.json sends {aggregate: 1}, which names no collection; dispatch resolved the namespace after taking the catalog lock and bailed with a plain return, holding it shared forever. A leaked shared lock is invisible to readers, so the server stayed perfectly responsive — an external prober got ok 15ms right through the hang — and only the next write that had to take the catalog exclusive to create a collection blocked. The failure therefore surfaced one file later, on a different connection, as a client-side timeout with nothing pointing at its cause.

Two lessons for using this runner:

  • A healthy-looking server does not exonerate the engine. Probe with the operation that is actually stuck, not with ping.

  • The driver's own command log is the fastest way in. It showed an insert sitting for exactly socketTimeoutMS against an idle engine, which is what turned a week-long-looking mystery into a five-line fix:

    MONGODB_LOG_COMMAND=debug MONGODB_LOG_PATH=stderr \
      node tests/spec/run.js --skip 68 --limit 2 2>drv.log
    

--skip/--limit exist for exactly this: the collapse reduced to a reproducible two-file window, which is what made it tractable.

Still worth recording the baseline on an otherwise idle machine, and do not tighten --op-timeout-ms to make a run finish sooner — a tight budget turns load into apparent engine failures, which is how I misdiagnosed this once already.