commands: the $group accumulators
Nine of them -- `$sum`, `$avg`, `$min`, `$max`, `$first`, `$last`, `$push`,
`$addToSet`, `$count` -- and the corpus goes 9 pass / 10 fail to 18 pass /
1 fail. The one left is the compound `_id`, which needs the expression
evaluator.
Landed before that evaluator, against the tier order the design review set
out, and the corpus is why: every one of these takes a single value per
document, a path or a constant, so nine of its ten failures turned out to be
reachable without one. `classify_expr` already produced exactly that value.
What the recording caught, which is the argument for measuring expectations
rather than writing them:
- `$avg` over a group with no numeric value is **null**, not `0`. A divisor
that counted documents rather than numbers would pass every test anybody
would think to write by hand, and be wrong on the one group that matters.
- `$min`/`$max` compare across types in canonical BSON order, so the maximum
of `30`, `7` and `"not a number"` is the string.
- `$push` skips an absent field but would push an explicit null, so "resolved
to nothing" and "resolved to null" cannot be the same value internally --
which is why the accumulators take `?bson.Value` and not `.null`.
- `$first`/`$last` follow input order, including when the value is absent:
`$last` of a missing field is null, not the last present one.
`AccState` is one struct rather than a union: the fields are small and every
site already switches on the kind, so a union would add a tag test where a
switch was going to be anyway. Its arrays are the gpa's, the values inside them
the reply arena's -- they outlive the group and travel with the documents.
`numeric_value` is the int32-or-double narrowing MongoDB reports, shared now
between the accumulators and `cmd_aggregate`'s count fast path. It was written
twice before; a divergence between them would make `countDocuments` disagree
with the pipeline it is a shortcut for.
`$avg` and `$push` came out of the Tier 0 refusal test, replaced by
`$stdDevPop` and `$mergeObjects`. The refusal is a property of what is missing
rather than of a list, and the test should read that way.
191/191 unit tests in ReleaseFast and ReleaseSafe, 83/83 fuzz, e2e 49, e2e3 16,
e2e4 17, e2e6 72, e2e7 86, crud corpus unchanged at 201/90/196.
This commit is contained in:
@@ -58,12 +58,15 @@ answer measures the version gap, not the engine.
|
||||
|
||||
## Where it stands
|
||||
|
||||
Recorded against mongod 8.3.7, run against the M2 tip:
|
||||
Recorded against mongod 8.3.7. At the M2 tip it read 9 pass / 10 fail; with the
|
||||
accumulators in:
|
||||
|
||||
```
|
||||
group-accumulators.json 9 pass 10 fail 0 skip
|
||||
group-accumulators.json 18 pass 1 fail 0 skip
|
||||
```
|
||||
|
||||
The nine include the four refusals M2 added, which answer with mongod's own
|
||||
codes. The ten are M2.5's work: `$avg`, `$min`, `$max`, `$first`, `$last`,
|
||||
`$push`, `$addToSet`, `$count`, and a compound `_id`.
|
||||
The one that remains is the compound `_id`, which needs the expression
|
||||
evaluator and is the next tier. The corpus found its first real disagreement on
|
||||
the way there: `$avg` over a group with no numeric value is `null`, not `0`,
|
||||
and a divisor that counted documents rather than numbers would have passed
|
||||
every test anybody would think to write by hand.
|
||||
|
||||
Reference in New Issue
Block a user