tests/spec: MongoDB spec-test runner and the M0 scorecard

PLAN D2 makes the official specification suites the gate for command semantics;
D7.6 asks for the harness to exist at M0 with a recorded baseline. This is that
harness, pinned on both sides -- mongodb/specifications @ 615e0f9 and
mongodb@7.5.0 -- because a scorecard is only comparable across milestones if a
delta cannot be an upstream test change.

It implements the unified format's Evaluating Matches algorithm as written,
including the two rules that decide whether a pass is earned: extra keys are
tolerated only in a root document, and numeric types compare flexibly. Anything
unimplemented is a SKIP with a reason, never a pass, and the one assertion class
not yet checked -- expectEvents, i.e. command monitoring -- is disclosed at the
top of the scorecard so `pass` reads as an upper bound.

First honest run: 131 pass, 161 fail, 195 skip over 175 files, zero timeouts.

Getting there took four attempts, and the failures are documented in the README
because each would have shipped a scorecard claiming a compatibility gap that
did not exist. Two were genuine leaks in this runner (clients left open when a
case timed out; clients registered for cleanup only after `await connect()`,
plus abandoned cases still creating more). The third I misdiagnosed as machine
load. The fourth attempt found the real cause: a leaked catalog lock in the
engine, fixed separately, which alone accounts for the jump from 45 passes to
131.

So the runner carries its own guards: per-operation CSOT timeouts so work is
never abandoned, an active-handle census per file, an end-of-run tripwire for
stray timers, a hard stop if the server dies rather than emitting hundreds of
misleading ECONNREFUSED failures, and --skip/--limit for bisecting a run whose
failures depend on position. The README states the rule plainly -- a long
unbroken tail of timeouts is a harness bug until proven otherwise -- and the two
commands that settle it.

Also fixes bench-run.sh, which copied its report over bench-latest.txt
unconditionally, including after a run that only warned -- so a degraded run
could silently replace the baseline that PLAN D7.5 makes a milestone gate.
This commit is contained in:
2026-08-03 18:58:01 +03:00
parent e2c25a986b
commit 90de7820da
7 changed files with 1487 additions and 6 deletions

View File

@@ -57,6 +57,7 @@ for C in $CLIENTS; do
MD_D=$(echo "$MD" | sed -E 's/.*\t([0-9.]+) docs\/s.*/\1/')
if [ -z "$MFDB_D" ] || [ -z "$MD_D" ]; then
echo "WARNING: concurrent run at $C clients produced no result (mfdb='$MFDB' md='$MD')" >&2
DEGRADED=1
fi
RATIO=$(node -e "const m=Number('$MFDB_D'),d=Number('$MD_D');console.log(d>0?(m/d).toFixed(1)+'x':'—')")
echo "clients $C $MFDB_D $MD_D $RATIO" | tee -a "$CC"
@@ -148,5 +149,17 @@ if [ -f "$LATEST" ]; then
else
echo "(no previous run to diff — this is the baseline)"
fi
# bench-latest.txt is the baseline the next run diffs against, and PLAN D7.5
# makes "no phase8 row regresses" a milestone gate -- so a run that only
# half-produced its numbers must not become the thing we compare to. This used
# to be an unconditional copy, which meant a degraded run silently replaced the
# baseline it was supposed to be measured against, and the missing rows then
# read as "no change" forever.
if [ "${DEGRADED:-0}" = "1" ]; then
echo >&2
echo "NOT updating $LATEST: this run was degraded (see WARNING above)." >&2
echo "The report is still at $REPORT if you want it." >&2
exit 1
fi
cp "$REPORT" "$LATEST"
echo; echo "latest: $LATEST"