index: a leaf record's payload becomes the document's slab offset

PLAN amendment A3. The B+tree leaf had nowhere to put a document's slab
offset -- `Slot.extra` is the payload length for a leaf and the child node id
for an internal separator -- which is what blocks the `_id_` tree from becoming
the primary lookup once the docs hashmap goes away.

A leaf record is now `key ++ offset_le`, so `extra` is always 8 and every
byte-accounting site (fits, record_cost, slot_cost, balanced_cut,
repack_keep_prefix) is untouched. Records get *smaller*: an ObjectId `_id_`
record goes from 26 bytes to 21.

`Entry.id` is deleted rather than re-owned. Every entry one document
contributes shares one document, so which document it is belongs on the call
that commits the entries -- which also makes it impossible to confuse the
offset a replace is removing with the one it is inserting. The old field
aliased the docs map's key and was only safe because removal happened at the
one chokepoint where a document dies; that constraint is gone.

Done for secondary indexes too, not just `_id_`. That deletes the per-candidate
`coll.docs.get(id)` in scan_sorted outright rather than replacing it with an
`_id_` descent, and it is free on the write path because a replace already
removes and reinserts every entry in every index.

Consequences worth knowing:

- lookup_eq/lookup_range/Plan.search yield u64. Those are values, immune to the
  tree mutation that invalidated the id slices they used to hand back -- which
  is why ttl_sweep_coll can drop the dupe-and-free dance it needed to survive
  `remove` freeing the key its entries pointed at.
- One safety net is gone. A stale entry used to be swallowed by
  `docs.get(id) orelse continue`; now it resolves to superseded-but-parseable
  bytes the re-applied filter might accept. That trades an invisible
  under-approximation for a visible wrong answer, which is the better failure
  to have, but it is a trade.
- A checkpoint may never renumber slab offsets (already recorded in PLAN §4):
  every index leaf now holds a physical one.

`zig build fuzz` earned its keep immediately -- it caught the API break in all
four B+tree harnesses, which `zig build test` cannot see.

Benchmarks A/B'd at 256m on one harness, before and after: all rows flat.
updateMany and deleteOne+insertOne first looked 10-13% slower, which three
repeat runs showed to be single-sample noise (0.70/0.71/0.70 against 0.70).
This commit is contained in:
2026-08-03 19:23:17 +03:00
parent 90de7820da
commit 491a4d0a6a
7 changed files with 320 additions and 275 deletions

View File

@@ -753,18 +753,22 @@ fn scan_sorted(
var n: usize = 0;
// Index plan (the implicit _id_ index first, then the secondaries):
// candidates in index order, re-filtered. The returned ids alias the
// docs map keys, valid under the read lock.
// candidates in index order, re-filtered. A candidate *is* a slab offset
// now, so the map lookup that used to translate an id into one is gone --
// and so is the accidental safety net it provided: a stale entry used to be
// dropped silently by `orelse continue`, where now it resolves to
// superseded-but-parseable bytes that the re-applied filter might accept.
// Loud beats silent: a wrong answer a test can see beats a missing
// candidate nothing can.
if (try index.plan(ctx.gpa, &coll.id_index, coll.indexes.items, filter, sort)) |p| {
var plan = p;
defer plan.deinit(ctx.gpa);
var ids: std.ArrayListUnmanaged([]const u8) = .empty;
defer ids.deinit(ctx.gpa);
try plan.search(ctx.gpa, &ids);
var offs: std.ArrayListUnmanaged(u64) = .empty;
defer offs.deinit(ctx.gpa);
try plan.search(ctx.gpa, &offs);
if (sorted) |flag| flag.* = plan.provides_sort;
if (plan.provides_sort) lim = limit;
for (ids.items) |id| {
const off = coll.docs.get(id) orelse continue;
for (offs.items) |off| {
if (!try query.matches_bytes(ctx.gpa, filter, coll.doc_bytes(off))) continue;
if (out) |list| try list.append(ctx.gpa, off);
n += 1;