tests/spec: MongoDB spec-test runner and the M0 scorecard
PLAN D2 makes the official specification suites the gate for command semantics; D7.6 asks for the harness to exist at M0 with a recorded baseline. This is that harness, pinned on both sides -- mongodb/specifications @ 615e0f9 and mongodb@7.5.0 -- because a scorecard is only comparable across milestones if a delta cannot be an upstream test change. It implements the unified format's Evaluating Matches algorithm as written, including the two rules that decide whether a pass is earned: extra keys are tolerated only in a root document, and numeric types compare flexibly. Anything unimplemented is a SKIP with a reason, never a pass, and the one assertion class not yet checked -- expectEvents, i.e. command monitoring -- is disclosed at the top of the scorecard so `pass` reads as an upper bound. First honest run: 131 pass, 161 fail, 195 skip over 175 files, zero timeouts. Getting there took four attempts, and the failures are documented in the README because each would have shipped a scorecard claiming a compatibility gap that did not exist. Two were genuine leaks in this runner (clients left open when a case timed out; clients registered for cleanup only after `await connect()`, plus abandoned cases still creating more). The third I misdiagnosed as machine load. The fourth attempt found the real cause: a leaked catalog lock in the engine, fixed separately, which alone accounts for the jump from 45 passes to 131. So the runner carries its own guards: per-operation CSOT timeouts so work is never abandoned, an active-handle census per file, an end-of-run tripwire for stray timers, a hard stop if the server dies rather than emitting hundreds of misleading ECONNREFUSED failures, and --skip/--limit for bisecting a run whose failures depend on position. The README states the rule plainly -- a long unbroken tail of timeouts is a harness bug until proven otherwise -- and the two commands that settle it. Also fixes bench-run.sh, which copied its report over bench-latest.txt unconditionally, including after a run that only warned -- so a degraded run could silently replace the baseline that PLAN D7.5 makes a milestone gate.
This commit is contained in:
@@ -57,6 +57,7 @@ for C in $CLIENTS; do
|
||||
MD_D=$(echo "$MD" | sed -E 's/.*\t([0-9.]+) docs\/s.*/\1/')
|
||||
if [ -z "$MFDB_D" ] || [ -z "$MD_D" ]; then
|
||||
echo "WARNING: concurrent run at $C clients produced no result (mfdb='$MFDB' md='$MD')" >&2
|
||||
DEGRADED=1
|
||||
fi
|
||||
RATIO=$(node -e "const m=Number('$MFDB_D'),d=Number('$MD_D');console.log(d>0?(m/d).toFixed(1)+'x':'—')")
|
||||
echo "clients $C $MFDB_D $MD_D $RATIO" | tee -a "$CC"
|
||||
@@ -148,5 +149,17 @@ if [ -f "$LATEST" ]; then
|
||||
else
|
||||
echo "(no previous run to diff — this is the baseline)"
|
||||
fi
|
||||
# bench-latest.txt is the baseline the next run diffs against, and PLAN D7.5
|
||||
# makes "no phase8 row regresses" a milestone gate -- so a run that only
|
||||
# half-produced its numbers must not become the thing we compare to. This used
|
||||
# to be an unconditional copy, which meant a degraded run silently replaced the
|
||||
# baseline it was supposed to be measured against, and the missing rows then
|
||||
# read as "no change" forever.
|
||||
if [ "${DEGRADED:-0}" = "1" ]; then
|
||||
echo >&2
|
||||
echo "NOT updating $LATEST: this run was degraded (see WARNING above)." >&2
|
||||
echo "The report is still at $REPORT if you want it." >&2
|
||||
exit 1
|
||||
fi
|
||||
cp "$REPORT" "$LATEST"
|
||||
echo; echo "latest: $LATEST"
|
||||
|
||||
Reference in New Issue
Block a user