PLAN D7's six items, with the numbers and the command that reproduces each in
tests/e2e/results/m0-gates.txt. Unit tests green in both optimize modes, the
whole e2e matrix green, the spec scorecard byte-identical at 131/161/195, and
the large smoke run at the scale D7.3 asked for:
21.47 GB collection (1,310,720 x 16 KiB)
data file 21.75 GB (+1.3% over the documents)
log after the load 2.5 MB (checkpoints reclaim it)
kill -9 then reopen 0.5 s (0.5 s at 4 GB too -- flat)
RSS after reopen 237 MB (1.1% of the data)
count after restart 1,310,720 last document byte-intact
acked writes after kill 200/200
That is the milestone's claim, measured: an open costs the working set rather
than the size of the database. Before M0 the same measurement was 523 MB
resident for a 512 MB database, because recovering each document's `_id` meant
reading every document at open.
Two gates need reading rather than a tick, and m0-gates.txt says so where a
reader would otherwise take a tick for granted.
The churn gate settles at 1.65x live data (delete-heavy) to 2.47x
(update-heavy), flat, above the ~1.3x amendment A2 hoped for. Rebuild-only
reclamation cannot reach that: it needs a whole second copy of the live data
before the first can be freed. The gate existed to decide whether doc-level
free lists are needed after M0, and that is the answer.
Benchmark parity holds for every read and latency row inside the run-to-run
spread, and bulk insert regresses 24% (732 -> 555 MB/s), reproducibly across
three runs. Risk 1 as written: document bytes now reach the disk uncompressed
on top of the LZ4 log. createIndex improves 62% from the same change.
Three measurement bugs fixed while running the gates, because each would have
put a false number in the README:
- `compare-run.sh` measured "db on disk" as `du` of the log alone against
`du` of mongod's whole dbpath. It reported 20 MB for a 1 GB collection --
the documents had moved to <db>.data. Honest figure, measured: 914 MB of
allocated blocks against mongod's compressed 85 MB.
- `big.js` counted "compaction events" as "the log shrank", which is a
*checkpoint* now. It claimed 12 compaction rewrites during a pure insert
load, which has no garbage to compact.
- `big.js` labelled peak RSS "in-memory engine: docs live in RAM" and its
summary said the collection was held "fully in RAM". Both were true of the
engine this milestone replaced.
README: the storage section described an all-in-RAM engine; the comparison
table mixed one old run's body with three new rows; and `findOne({_id})` was
documented as a full scan for integer ids, which the ordered `_id_` index made
false (2 ms against 55 s for a scan of the same 21.5 GB collection). The table
is now best-of-three for both servers, with the measured variance stated, since
two runs of the same binary moved the sub-10 ms rows by 27-51%.
399 lines
22 KiB
Markdown
399 lines
22 KiB
Markdown
# MultiforaDB
|
||
|
||
A lightweight, embedded MongoDB-compatible document database written in
|
||
Zig 0.16. Like SQLite, it stores everything in a single file; unlike SQLite,
|
||
it speaks the MongoDB wire protocol, so real clients — `mongosh`, the Node.js
|
||
driver, PyMongo — connect over TCP and just work.
|
||
|
||
## Forward plan
|
||
|
||
The direction from this MVP — a full-fledged embedded, tens-of-GB,
|
||
maximally MongoDB-compatible database — its decision record, milestones
|
||
and gates live in [PLAN.md](PLAN.md). Milestone 0 (mmap + WAL storage
|
||
foundation) has landed; its measured gate results are in
|
||
[`tests/e2e/results/m0-gates.txt`](tests/e2e/results/m0-gates.txt).
|
||
Milestone 1 (cursors, and the doc-level free list the churn gate showed is
|
||
needed) is next.
|
||
|
||
## Quick start
|
||
|
||
```sh
|
||
zig build # build the server
|
||
zig build test # run the unit test suite
|
||
|
||
zig-out/bin/multiforadb --port 27017 --db data.log --compact-threshold 256m
|
||
|
||
# in another terminal:
|
||
mongosh --port 27017
|
||
> db.users.insertOne({name: "alice", age: 30})
|
||
> db.users.find({age: {$gt: 25}}).toArray()
|
||
> db.users.updateOne({name: "alice"}, {$set: {vip: true}})
|
||
> db.users.deleteOne({name: "bob"})
|
||
> db.sessions.createIndex({expireAt: 1}, {expireAfterSeconds: 3600})
|
||
```
|
||
|
||
## Features
|
||
|
||
- **Wire protocol**: OP_MSG (2013) plus legacy OP_QUERY/OP_REPLY (2004/2001)
|
||
for the driver handshake; hello/isMaster with `maxWireVersion: 8`, so
|
||
modern drivers (Node, Python, mongosh) connect without workarounds.
|
||
- **BSON**: full parse/serialize round-trip for all common types
|
||
(including binary, regex, timestamps, ObjectId), canonical MongoDB
|
||
comparison order for sorting and range queries.
|
||
- **CRUD**: `insert`, `find` (filter, sort, skip/limit, projection),
|
||
`update` (multi/upsert), `delete`, `findAndModify`, `count`,
|
||
`aggregate` (`$match`, `$sort`, `$skip`, `$limit`, `$project`, `$count`,
|
||
`$group` with `$sum`), plus `create`/`drop`/`listCollections`/
|
||
`listDatabases`/`dropDatabase`.
|
||
- **Query operators**: `$eq` `$ne` `$gt` `$gte` `$lt` `$lte` `$in` `$nin`
|
||
`$exists` `$regex` (hand-rolled engine: anchors, `.`, `* + ?`, character
|
||
classes, groups, alternation, `i`/`s` options) `$not` `$and` `$or` `$nor`
|
||
`$size` `$all` `$elemMatch`, with dot paths and array multikey semantics.
|
||
- **Secondary indexes**: `createIndex`/`listIndexes`/`dropIndex` via the
|
||
three driver commands, single-field and compound, with `unique`,
|
||
`sparse` and `expireAfterSeconds` (TTL) options, persisted in the log
|
||
and rebuilt on open (compaction re-emits them). A background sweeper
|
||
expires TTL-indexed documents through the ordinary logged write path.
|
||
The query planner turns equality / `$in` / range
|
||
predicates into index lookups across `find`, `count`, `update`,
|
||
`delete`, `findAndModify`, and a leading `$match` in `aggregate`; every
|
||
candidate is re-checked against the full filter, so an index that
|
||
over-approximates is merely slow, never wrong.
|
||
- **Update operators**: `$set` `$unset` `$inc` `$push` (`$each`) `$pull`
|
||
`$rename`, with dot-path creation (including array indices).
|
||
- **Storage**: append-only record log, LZ4-compressed in 256 KiB blocks
|
||
(XxHash3-checked, `fsync` per write, torn-tail tolerant: a crash
|
||
mid-append truncates cleanly, interior corruption is rejected) with
|
||
in-memory indexes rebuilt on open and automatic compaction (rewrite +
|
||
atomic rename when the log grows past `--compact-threshold`, default
|
||
16 MB). Killed mid-write (`kill -9`), the database recovers all
|
||
committed writes; the log and compaction both work with relative or
|
||
absolute `--db` paths. Records up to the announced 16 MB
|
||
`maxBsonObjectSize` replay correctly.
|
||
- **Concurrency**: a writer-preferring read/write lock splits command
|
||
execution — reads (`find`, `count`, `aggregate`, `list*`) run concurrently
|
||
across connections, writes (CRUD, DDL) are exclusive and totally ordered,
|
||
and handshake/no-op commands run lock-free. The log append + `fsync` still
|
||
happen under the write lock, so the crash guarantees are unchanged. Fine
|
||
for light workloads.
|
||
|
||
## Layout
|
||
|
||
```
|
||
src/
|
||
bson.zig BSON parse/serialize, ObjectId, canonical comparison order
|
||
wire.zig OP_MSG/OP_QUERY framing, message + reply builders
|
||
commands.zig command dispatch (hello, CRUD, aggregate, admin, indexes)
|
||
server.zig TCP accept loop, per-connection handlers, TTL sweep monitor
|
||
db.zig in-memory engine: db → collection → _id → document maps
|
||
storage.zig append-only log: records, replay, CRC validation
|
||
query.zig filter matcher, regex engine, sort, projection
|
||
index.zig secondary indexes: entries, search, query planner
|
||
update.zig update operators with dot-path navigation
|
||
main.zig CLI: --port, --bind, --db, --ttl-sweep-secs, --compact-threshold
|
||
```
|
||
|
||
## Indexes
|
||
|
||
`collection.createIndex({field: 1})` works against every driver; the index
|
||
is persisted in the log, survives restarts and compaction, and is used by
|
||
the query planner to narrow scans.
|
||
|
||
- **Key patterns**: single-field and compound (up to 32 fields), each key
|
||
`1` or `-1`. Descending order is metadata (entries are always stored
|
||
value-ascending); the default index name is MongoDB's `a_1_b_-1`.
|
||
`createIndex({_id: 1})` is an idempotent no-op — the docs map is the
|
||
`_id_` index — and `dropIndex("_id_")` errors.
|
||
- **Options**: `unique` (a conflicting write fails with E11000 naming the
|
||
index; per-document entries are deduped first, so `{a: [1,1]}` is legal)
|
||
and `sparse` (documents missing an indexed field are skipped).
|
||
- **TTL**: `createIndex({expireAt: 1}, {expireAfterSeconds: 60})` deletes a
|
||
document once its indexed date is that many seconds old. A background
|
||
sweeper runs every `--ttl-sweep-secs` seconds (default 60, `0` disables
|
||
it) and deletes through the ordinary write path, so each expiry is logged
|
||
and fsynced and holds across a restart. As in MongoDB the option is
|
||
single-field only (a compound key is `CannotCreateIndex`, code 67),
|
||
`expireAfterSeconds` must be a whole number in `[0, 2147483647]` (`0`
|
||
means "expire at the stored instant"), a non-date value at the path never
|
||
expires, an array of dates expires on its earliest member, and expiry is
|
||
coarse: a document stays visible until the next sweep. Re-creating an
|
||
index with a different expiry is `IndexOptionsConflict` (85) and an
|
||
expiry on `{_id: 1}` is `InvalidIndexSpecificationOption` (197), both as
|
||
MongoDB has them.
|
||
- **Multikey**: an array at an indexed path is indexed as a whole *and*
|
||
element-wise, mirroring the query matcher exactly, so both
|
||
`{tags: "a"}` and `{tags: ["a","b"]}` hit the index. A compound index
|
||
over two array paths rejects the document with MongoDB's "cannot index
|
||
parallel arrays".
|
||
- **Planner**: picks the index covering the longest leading run of
|
||
equality/`$in` predicates (cartesian product capped at 100 lookups),
|
||
optionally with a range on the next key. Ranges with both bounds fall
|
||
back to a scan on multikey indexes (a doc with `{a: [1,2]}` can satisfy
|
||
`{a: {$gt: 5, $lt: 25}}` across two entries), and sparse indexes are
|
||
never used for `null`-valued predicates. The `_id_` fast path resolves
|
||
`{_id: ...}` through the docs map unless the value's compare class is
|
||
serialization-ambiguous (int32 1, int64 1, double 1.0 compare equal but
|
||
hash differently — those fall back to a scan, as do string/symbol/code).
|
||
|
||
v1 limits: no hashed/text/geo/partial indexes, and entry *insert* is O(n)
|
||
(a sorted array memmoves the tail) — fine for a light database, with a
|
||
B-tree as the follow-up. Removal is no longer a scan: entry generation is
|
||
a pure function of the document, so the entries to drop are regenerated
|
||
and found by binary search. A TTL sweep
|
||
walks every entry of every TTL index and holds the write lock for the
|
||
whole pass, so the interval is the tuning knob: expiry is never more
|
||
precise than `--ttl-sweep-secs`, and a very large TTL index wants a
|
||
longer one.
|
||
|
||
## Not (yet) implemented
|
||
|
||
- Authentication (SCRAM) — run without credentials
|
||
- Real cursors (all results are returned in one batch, cursor id 0)
|
||
- Transactions, change streams, replicasets
|
||
- Compression (OP_COMPRESSED)
|
||
- `collMod`, so an index's `expireAfterSeconds` cannot be changed in
|
||
place — drop the index and re-create it with the new expiry
|
||
- `dropCollection`/`dropDatabase` write no log record, so a dropped
|
||
collection (and its index definitions) resurrect on restart
|
||
|
||
## Working with large collections
|
||
|
||
Documents, B+tree pages and overflow records live in an mmap'd data file
|
||
(`<db>.data`), with the append-only log as the write-ahead log in front of it.
|
||
Every write command is logged with `fsync` before it is acknowledged (one sync
|
||
per command via group commit — a 500-doc `insertMany` syncs once, not 500
|
||
times); a periodic checkpoint publishes the data file and truncates the log.
|
||
So resident memory is the working set rather than the size of the database, and
|
||
an open does not replay everything ever written. Measured by the
|
||
`tests/e2e/big.js` harness on a 12-core/32 GB Mac:
|
||
|
||
| 21.5 GB collection (1.3M × 16 KiB) | |
|
||
|---|---|
|
||
| data file | 21.75 GB (+1.3% over the documents) |
|
||
| log after the load | 2.5 MB — checkpoints reclaim it |
|
||
| `kill -9` then reopen | 0.5 s (also 0.5 s at 4 GB) |
|
||
| resident after reopen | 237 MB — 1.1% of the data |
|
||
| documents after restart | all 1,310,720, last one byte-intact |
|
||
| acked writes surviving `kill -9` | 200/200 |
|
||
|
||
Notes on the cost side, from the same run:
|
||
|
||
- **A bulk load still touches everything it writes.** Peak resident during the
|
||
20 GB load was 18.5 GB: writing 21 GB of pages dirties 21 GB of pages, and
|
||
the kernel keeps them until it wants the memory back. The mmap win is in
|
||
reopen and steady-state reads, not in bulk ingest.
|
||
- **Bulk insert costs about a quarter of its old throughput** (732 → 555 MB/s
|
||
at 1 GB), because document bytes now reach the disk uncompressed in the data
|
||
file on top of the compressed log. This was the anticipated trade for the
|
||
rows above; see `m0-gates.txt` for the untried mitigations.
|
||
- **A cold full scan reads the whole collection from disk** — ~55 s for 21 GB,
|
||
about 390 MB/s. Index the fields you filter on; `countDocuments({})` with no
|
||
filter is a full scan by definition.
|
||
- **Churn is bounded but not tight.** Under sustained rewriting the data file
|
||
settles at 1.65× (delete-heavy) to 2.47× (update-heavy) the live data and
|
||
stays there. Reclamation is by whole-collection rebuild, which needs a second
|
||
copy of the live data before it can free the first; a doc-level free list is
|
||
the M1 fix.
|
||
|
||
- **Build in ReleaseFast** — `zig build` defaults to it. A Debug server is
|
||
10-200x slower on every path (the matcher alone was 70 µs/doc in Debug
|
||
vs 0.4 µs in ReleaseFast), which dwarfed every other difference in the
|
||
MongoDB comparison below.
|
||
- **Compaction no longer needs tuning for bulk loads.** It triggers on the
|
||
share of the log that is garbage rather than on bytes appended, so a pure
|
||
insert workload — which has no garbage — is never rewritten, and a
|
||
rewrite-heavy one is reclaimed once about a fifth of the log is dead,
|
||
keeping the file near 1.25x the live data. `--compact-threshold` is now
|
||
only a floor below which small logs are left alone. (It used to fire
|
||
every 16 MB regardless, rewriting the whole log each time: quadratic
|
||
total traffic, and the reason bulk loads needed a raised threshold.)
|
||
- **`findOne({_id})` is an index descent for every `_id` type.** It used to
|
||
be O(1) for ObjectIds and a *full scan* for integer, int64 and double ids,
|
||
which compare equal but hashed differently. The hash map is gone: `_id_` is
|
||
an ordered B+tree over the canonical key encoding, so all of those are one
|
||
descent. Measured on the 21.5 GB collection with integer ids:
|
||
`findOne({_id})` 2 ms, against 55 s for a scan of the same collection.
|
||
One consequence to know about: because the encoding is canonical, `1`
|
||
(int32), `1` (int64) and `1.0` (double) are now the *same* `_id` — which
|
||
matches MongoDB, and which a database written by an older build will warn
|
||
loudly about on first open if it holds two such documents.
|
||
- **Secondary-index entry insert is O(n)** (sorted array — see v1 limits
|
||
above), so inserting into a collection that already has an index is
|
||
quadratic. Building an index over existing data is not: entries are
|
||
appended unsorted and ordered once. Still cheapest to create indexes
|
||
after the load.
|
||
|
||
## Performance vs MongoDB
|
||
|
||
`tests/e2e/compare-run.sh` runs the same driver workload (1 GB, 65,536 ×
|
||
16 KB docs, every write durable — MultiforaDB fsyncs per command, mongod
|
||
runs with `j: true`) against each server and prints a side-by-side table.
|
||
With the ReleaseFast default build (MongoDB 8.3.7 on the same Mac):
|
||
|
||
| benchmark | MultiforaDB | mongodb | winner |
|
||
|---|---|---|---|
|
||
| insertOne (sequential) ×200 | 0.20 ms | 4.4 ms | **MultiforaDB ×22** |
|
||
| bulk insert (insertMany) | 555 MB/s | 739 MB/s | mongodb ×1.3 |
|
||
| createIndex({k: 1}) | 27.9 ms | 70 ms | **MultiforaDB ×3** |
|
||
| countDocuments({}) | 2.6 ms | 11.3 ms | **MultiforaDB ×4** |
|
||
| findOne({_id}) | 0.60 ms | 0.61 ms | parity |
|
||
| findOne indexed | 0.54 ms | 1.1 ms | **MultiforaDB ×2** |
|
||
| range-scan count | 11.2 ms | 12.5 ms | **MultiforaDB** |
|
||
| sort + limit(20) on `_id` | 1.5 ms | 2.0 ms | **MultiforaDB** |
|
||
| projection + limit(1000) | 3.6 ms | 4.3 ms | **MultiforaDB** |
|
||
| aggregate $group | 7.4 ms | 13.4 ms | **MultiforaDB ×2** |
|
||
| updateOne({_id}) ×50 | 0.15 ms | 0.21 ms | **MultiforaDB** |
|
||
| updateMany (65 docs) | 1.9 ms | 5.9 ms | **MultiforaDB ×3** |
|
||
| deleteOne + insert | 0.63 ms | 4.9 ms | **MultiforaDB ×8** |
|
||
| concurrent durable writes, 32 clients | 32,817/s | 3,765/s | **MultiforaDB ×9** |
|
||
| server RSS after the load | 1.06 GB | 1.18 GB | MultiforaDB |
|
||
| kill -9 → reopen | 0.3 s | 1.3 s | **MultiforaDB ×4** |
|
||
| db on disk | 914 MB | 85 MB | **mongodb ×11** |
|
||
|
||
Each cell is the best of three runs of the same suite, for both servers. That
|
||
is not fussiness: two consecutive runs of the *same* binary moved the sub-10 ms
|
||
rows by 27–51% on this machine, so a single run's ratios say more about the
|
||
minute they were taken in than about either database. Treat differences under
|
||
about 1.5× as noise.
|
||
|
||
Two rows in that table changed direction with the mmap foundation and are
|
||
worth being explicit about.
|
||
|
||
**`db on disk`** was 97 MB against mongod's 91 MB when the log was the only
|
||
copy of the data and LZ4 compressed it. The data file does not compress
|
||
documents: 65,536 × 16 KiB documents now occupy 914 MB of allocated blocks
|
||
(`du`; `ls` shows 1.09 GB, the difference being the sparse tail the file is
|
||
grown into) against mongod's compressed 85 MB. Per-page or per-extent
|
||
compression is the fix, and it is not in M0. If you are comparing against a
|
||
report from before this was written, note that the row used to measure the log
|
||
file alone — which said 20 MB for a 1 GB collection, because the documents had
|
||
moved to `<db>.data`. Fixed in `compare-run.sh`.
|
||
|
||
**`server RSS`** is measured right after writing the whole dataset, so it
|
||
reflects a bulk load having dirtied every page it wrote, not steady state. The
|
||
number that speaks to the architecture is resident memory *after a reopen*:
|
||
237 MB for a 21.5 GB collection (see the section above). Before M0 the same
|
||
measurement was 523 MB for a 512 MB collection, because recovering each
|
||
document's `_id` read every document at open.
|
||
|
||
The range scan runs at parity because matching happens against the stored BSON
|
||
bytes directly, skipping fields by length, with no per-document arena and no
|
||
second Pair-tree copy. The log is still LZ4-compressed in 256 KiB blocks; since
|
||
it is now truncated at every checkpoint, its size no longer tracks the
|
||
database's.
|
||
|
||
Reproduce the whole thing — main suite, concurrency sweep and the meta rows —
|
||
with `bash tests/e2e/bench-run.sh 1g 16k "1 4 8 16 32"`, which writes a
|
||
timestamped report to `tests/e2e/results/` and diffs it against the last one.
|
||
The pre-tree baseline is in `tests/e2e/results/phase1.txt`; the runs with the
|
||
B+tree, ordered `_id` index, compressed log and byte storage (roadmap items
|
||
1–4) in `phase2.txt` through `phase5.txt`; the all-in-RAM engine this table's
|
||
predecessor measured in `phase8.txt`; and the mmap foundation's own gate
|
||
results, including what each of those rows cost or gained, in
|
||
`m0-gates.txt`.
|
||
|
||
### What is left (highest impact first)
|
||
|
||
Each is written up with its design decisions, ordering constraints and
|
||
traps in [ROADMAP.md](ROADMAP.md).
|
||
|
||
The remaining structure is the single-file log (appends and the commit
|
||
serialize on one log lock, though appends no longer hold the collection
|
||
locks), and the acknowledged-write fsync, which dominates sequential
|
||
per-client workloads. All five roadmap items are landed.
|
||
|
||
Done so far, with the measurement that drove each:
|
||
|
||
- **Record integrity hash CRC32 → XxHash3.** `std.hash.Crc32` is
|
||
table-driven and byte-at-a-time: 408 MB/s against XxHash3's 31 GB/s, or
|
||
38 µs versus 0.5 µs on a 16 KB document — about two thirds of the entire
|
||
bulk-insert cost. Insert 260 → 700 MB/s.
|
||
- **Compaction triggers on garbage, not on bytes written**, and syncs once
|
||
per rewrite instead of once per document. Bulk load at the default
|
||
threshold 41.6 → 703 MB/s.
|
||
- **Index builds append then sort once** instead of inserting into a sorted
|
||
array. `createIndex` over 65,536 documents 649 → 44 ms.
|
||
- **Index entries hold encoded byte keys**, so comparing them is a memcmp
|
||
rather than a walk over values in unrelated arenas.
|
||
- **A B+tree over the encoded keys** (roadmap item 1): fixed 4 KiB slotted
|
||
pages in a flat u32-addressed node array, an overflow slab for long
|
||
records, no rebalancing on delete, and bulk bottom-up packing. Entry
|
||
insertion and removal are a descent plus a leaf-local edit instead of a
|
||
tail memmove, so writes into an already-built index stopped being
|
||
quadratic. `updateMany` 17.3 → 1.6 ms (2.8x slower than MongoDB → 4x
|
||
faster); `createIndex` 62 → 51 ms.
|
||
- **An ordered `_id` index** (roadmap item 2): every collection carries an
|
||
implicit `_id_` index (kept out of the secondary list, so the listing,
|
||
drop and log-format surfaces are unchanged; rebuilt after replay like
|
||
the secondaries). Its encoded keys are canonical, so the old
|
||
serialization-guarded docs-map fast path is gone and integer/string
|
||
`_id` point lookups, `$in` and ranges hit the tree instead of a full
|
||
scan. `sort({_id: ...})` is now an index-ordered scan with an early stop:
|
||
`sort+limit(20)` 6.2 → 2.4 ms (parity with MongoDB).
|
||
- **A block-framed, LZ4-compressed log** (roadmap item 3): a file header
|
||
plus ~256 KiB blocks, each holding the existing record framing with the
|
||
integrity hash covering the stored bytes (so the decompressor only ever
|
||
sees input already proven intact). Records never straddle blocks; a
|
||
short read, impossible length or hash mismatch in the final block is a
|
||
torn tail (truncate cleanly), anywhere else is corruption. The hand-rolled
|
||
LZ4 codec runs at ~1.7 GB/s and falls back to raw per block when
|
||
compression does not help. `db on disk` 1025 → 97 MB — now smaller than
|
||
MongoDB's own compressed files.
|
||
- **Byte storage without per-document arenas** (roadmap item 4): documents
|
||
live as canonical BSON bytes in a segmented per-collection slab; the
|
||
docs map holds flat offsets (stable across segment growth, ≤ one segment
|
||
of slack). The matcher walks the bytes directly, skipping by length any
|
||
field the filter does not name (differential-tested against the tree
|
||
matcher on a corpus), and the scan/aggregate paths never materialize
|
||
stored documents; sort, projection, updates and index entry generation
|
||
use a borrowed spine into the slab. `server RSS` 1979 → 539 MB (2.4x
|
||
smaller than MongoDB); `range-scan` 22.5 → ~12 ms (parity, best run
|
||
faster); `proj` 4.1 → 3.4 ms.
|
||
- **Decomposed locks** (roadmap item 5): collections are heap-allocated;
|
||
a catalog rwlock guards the maps and each collection has its own rwlock
|
||
(catalog → collection → log ordering, one collection at a time for the
|
||
TTL sweep and compaction). Appends never fsync; a write command's
|
||
epilogue commits once with a leader/follower group commit, and
|
||
compaction snapshots collections without the log lock, retrying if a
|
||
writer appended mid-snapshot. Acknowledged writes are fsynced before
|
||
their reply; an unacknowledged write may vanish (ordinary `w:1, j:true`,
|
||
no longer "the log describes ≥ memory"). Concurrent durable-insert
|
||
throughput scales ~5.1k → 12.5k docs/s from 1 → 8 clients, ~14.8k at 32.
|
||
- **Entry removal is a binary search**, not a scan of the whole index.
|
||
`updateMany` 15.4 → 5.5 ms.
|
||
- **Top-k sort selection** and an allocation-free decorate pass, plus
|
||
**index-supplied ordering** when an index already holds candidates in the
|
||
requested order. `sort+limit(20)` 40 → 4.3 ms, or 1.0 ms on an indexed
|
||
field.
|
||
- **`limit` reaches the scan**, which used to materialize the whole
|
||
collection before slicing, and `countDocuments` is answered by counting
|
||
rather than by materializing and discarding every match.
|
||
- **Matching collects candidates on the stack**, resolves operators to an
|
||
enum once per filter field rather than by string per document, and reuses
|
||
one reply arena per connection.
|
||
|
||
Several real bugs surfaced while benchmarking:
|
||
|
||
- `plan_id` returned a pointer to a stack temporary (`&.{e}`) that dangled
|
||
after the frame returned — Debug tolerated it, ReleaseFast read garbage,
|
||
silently breaking every `findOne({_id: <ObjectId>})`. It now heap-copies
|
||
the lookup value and frees it.
|
||
- Multi-doc writes fsynced once per document; they now group-commit (one
|
||
fsync per command, same crash guarantees — verified by the kill -9
|
||
crash suites).
|
||
- Compaction fsynced once per live document, because the log it wrote into
|
||
never had deferred syncing enabled — 65,536 fsyncs to rewrite a 1 GB
|
||
collection.
|
||
- `remove` never checked the compaction threshold, so a delete-heavy
|
||
workload grew the log without bound.
|
||
|
||
|
||
## Code style
|
||
|
||
The project follows TigerBeetle's TigerStyle — see
|
||
[`docs/TIGER_STYLE.md`](docs/TIGER_STYLE.md) (binding reference) and the
|
||
"Code style" section of `AGENTS.md` for the project-specific rules and
|
||
deliberate deviations. Highlights: `zig fmt` clean, 100-column hard limit,
|
||
4-space indent, snake_case, functions under 70 lines, always-on
|
||
assertions via `src/assert.zig`.
|