Files
MultiforaDB/tests/spec/aggregate/README.md
A.Shakhmatov 1c72aa8938 tests/spec: record the expression evaluator's spec from mongod
27 cases, recorded before a line of the evaluator is written, which is the
whole point of having built the recorder first: the expectations come from
mongod 8.3.7 rather than from what the implementation is about to do.

The corpus reads 1 pass / 26 fail, and the one pass is the unknown-operator
refusal M2 already answers correctly. Expressions are exercised through
`$group` because `_id` and the accumulator arguments are the only expression
positions that exist until `$addFields` and `$project`'s computed fields land;
testing them anywhere else would be testing a stage that is not there.

Ten things the recording settled that guessing would have got wrong:

    $add over a missing field or null   null -- not an error, and not 0
    $add over a string                  error 7157723
    $divide by zero                     error 4848401
    $mod of -5 by 4                     -1, the dividend's sign
    $lt of a number and a string        true, canonical type order
    $and over -5                        truthy
    $not of a missing field             true
    $switch, no branch and no default   error 40069
    two operators in one expression     error 15983, *not* $group's 40238
    $subtract with one operand          error 16020

The six error codes are in `ErrorCode` already, so the evaluator's refusals and
its runtime failures have somewhere measured to land. The evaluator itself is
the next commit and is not started.

191/191 unit tests, crud corpus unchanged at 201/90/196.
2026-08-09 22:12:48 +03:00

94 lines
3.8 KiB
Markdown

# The aggregation corpus
`mongodb/specifications` has no aggregation suite. The thirteen
`aggregate-*.json` files this project runs come from `crud` and test the
aggregate *command* — cursors, read concern, the write stages, collation,
`let`. They touch stages barely and expressions not at all: `$lookup`,
`$unwind`, `$facet`, `$addFields` and `$replaceRoot` appear nowhere in the
pinned corpus. That is PLAN amendment A6, and this directory is its
consequence: M2.5 has to bring its own gate.
## The one rule
**Inputs are authored here; expectations are measured against a real mongod.**
A corpus we write is a corpus that can encode our own bugs as expectations, and
it would then agree with us forever. So `sources/*.json` holds documents and
pipelines and nothing else, and `record.js` asks mongod 8.3.7 what each pipeline
answers. It is the same discipline that corrected three assumptions in M1's
session work and every error code in M2 — the alternative, in both cases, would
have shipped.
## Running it
```sh
node tests/spec/run.js --suite-dir tests/spec/aggregate
```
The same runner as the crud corpus, pointed elsewhere. Sharing it is the point:
the entity model, the matchers, the skip accounting and `expectEvents` come for
free, and a second runner would drift from the first exactly where it mattered.
`--scorecard` is refused with `--suite-dir`: `tests/spec/scorecard.txt` is the
crud corpus's record and the milestones are compared against it.
## Re-recording
```sh
mongod --port 27099 --dbpath /tmp/mongo-corpus &
node tests/spec/aggregate/record.js --mongod-port 27099
```
Writes `<name>.json` for every `sources/<name>.json`. The generated files are
committed: they *are* the corpus, and regenerating them is how a disagreement
with mongod gets re-measured rather than argued about.
Two things to know when adding cases:
- **End a `$group` pipeline with a `$sort`.** Group output order is unspecified,
and a case that depended on it would fail for the wrong reason on either
server.
- **Errors record the code, not the message.** Message text is mongod's to
change between releases; a corpus that pinned it would break for the wrong
reason.
Leave out any case whose answer depends on a server newer than the 4.4 this
server reports — recording it from mongod 8.x and judging it against a 4.4
answer measures the version gap, not the engine.
## Where it stands
Recorded against mongod 8.3.7. At the M2 tip it read 9 pass / 10 fail; with the
accumulators in:
```
group-accumulators.json 18 pass 1 fail 0 skip
expressions.json 1 pass 26 fail 0 skip
```
`group-accumulators` found its first real disagreement on the way to 18:
`$avg` over a group with no numeric value is `null`, not `0`, and a divisor
that counted documents rather than numbers would have passed every test
anybody would think to write by hand.
`expressions.json` is the next tier's spec, recorded and not yet implemented --
its one pass is the unknown-operator refusal M2 already answers correctly.
Expressions are exercised through `$group`, because `_id` and the accumulator
arguments are the only expression positions that exist until `$addFields` and
`$project`'s computed fields land.
What recording it settled, none of which is guessable:
| | mongod |
|---|---|
| `$add` over a missing field or null | `null`, not an error and not `0` |
| `$add` over a string | error 7157723 |
| `$divide` by zero | error 4848401 |
| `$mod` of -5 by 4 | `-1` — the dividend's sign |
| `$lt` of a number and a string | `true` — canonical type order |
| `$and` over `-5` | truthy |
| `$not` of a missing field | `true` |
| `$switch` with no branch and no default | error 40069 |
| two operators in one expression document | error 15983, *not* `$group`'s 40238 |
| `$subtract` with one operand | error 16020 |