The log is now a 16-byte file header (magic, version, codec, block target) plus a sequence of blocks. Each block keeps the pre-existing record framing unchanged, so Engine.apply_record does not change; records never straddle blocks (appends accumulate in memory and the block seals at ~256 KiB). The block header's integrity hash covers the stored payload bytes exactly as they sit on disk, so the decompressor only ever sees input already proven intact. Torn tails stay distinguishable from interior corruption exactly as before: a short read, an impossible length, or a hash mismatch in the final block truncates cleanly (later appends overwrite the garbage); a hash mismatch anywhere else is error.InvalidLog. The codec is a hand-rolled LZ4 block compressor/decompressor (~1.7 GB/s measured) with a per-block codec byte falling back to raw when compression does not help; the header keeps raw legal so zstd can be swapped in later. Zig 0.16 ships zstd decompression only, and deflate would cap writes below the insert rate. Engine.compact goes through the same Log API (deferred sync, one commit) and compresses for free; sync() seals the pending block before fsyncing, so the acknowledged-write durability semantics are unchanged (an unsealed block holds only unacknowledged batch records). Measured (tests/e2e/results/phase4.txt): db on disk 1025 -> 97 MB, now smaller than MongoDB's own compressed files; bulk insert 816 -> 722 MB/s (the accepted compression cost); reopen unchanged at 0.8 s. Verified: unit suite in all three optimize modes (new LZ4 round-trip, corrupt-block, and torn-tail truncation tests), the crash pair, e2e6 (kill -9 mid-write), and two full benchmark runs.
307 lines
16 KiB
Markdown
307 lines
16 KiB
Markdown
# mongo-lite
|
||
|
||
A lightweight, embedded MongoDB-compatible document database written in
|
||
Zig 0.16. Like SQLite, it stores everything in a single file; unlike SQLite,
|
||
it speaks the MongoDB wire protocol, so real clients — `mongosh`, the Node.js
|
||
driver, PyMongo — connect over TCP and just work.
|
||
|
||
## Quick start
|
||
|
||
```sh
|
||
zig build # build the server
|
||
zig build test # run the unit test suite
|
||
|
||
zig-out/bin/mongo-lite --port 27017 --db data.log --compact-threshold 256m
|
||
|
||
# in another terminal:
|
||
mongosh --port 27017
|
||
> db.users.insertOne({name: "alice", age: 30})
|
||
> db.users.find({age: {$gt: 25}}).toArray()
|
||
> db.users.updateOne({name: "alice"}, {$set: {vip: true}})
|
||
> db.users.deleteOne({name: "bob"})
|
||
> db.sessions.createIndex({expireAt: 1}, {expireAfterSeconds: 3600})
|
||
```
|
||
|
||
## Features
|
||
|
||
- **Wire protocol**: OP_MSG (2013) plus legacy OP_QUERY/OP_REPLY (2004/2001)
|
||
for the driver handshake; hello/isMaster with `maxWireVersion: 8`, so
|
||
modern drivers (Node, Python, mongosh) connect without workarounds.
|
||
- **BSON**: full parse/serialize round-trip for all common types
|
||
(including binary, regex, timestamps, ObjectId), canonical MongoDB
|
||
comparison order for sorting and range queries.
|
||
- **CRUD**: `insert`, `find` (filter, sort, skip/limit, projection),
|
||
`update` (multi/upsert), `delete`, `findAndModify`, `count`,
|
||
`aggregate` (`$match`, `$sort`, `$skip`, `$limit`, `$project`, `$count`,
|
||
`$group` with `$sum`), plus `create`/`drop`/`listCollections`/
|
||
`listDatabases`/`dropDatabase`.
|
||
- **Query operators**: `$eq` `$ne` `$gt` `$gte` `$lt` `$lte` `$in` `$nin`
|
||
`$exists` `$regex` (hand-rolled engine: anchors, `.`, `* + ?`, character
|
||
classes, groups, alternation, `i`/`s` options) `$not` `$and` `$or` `$nor`
|
||
`$size` `$all` `$elemMatch`, with dot paths and array multikey semantics.
|
||
- **Secondary indexes**: `createIndex`/`listIndexes`/`dropIndex` via the
|
||
three driver commands, single-field and compound, with `unique`,
|
||
`sparse` and `expireAfterSeconds` (TTL) options, persisted in the log
|
||
and rebuilt on open (compaction re-emits them). A background sweeper
|
||
expires TTL-indexed documents through the ordinary logged write path.
|
||
The query planner turns equality / `$in` / range
|
||
predicates into index lookups across `find`, `count`, `update`,
|
||
`delete`, `findAndModify`, and a leading `$match` in `aggregate`; every
|
||
candidate is re-checked against the full filter, so an index that
|
||
over-approximates is merely slow, never wrong.
|
||
- **Update operators**: `$set` `$unset` `$inc` `$push` (`$each`) `$pull`
|
||
`$rename`, with dot-path creation (including array indices).
|
||
- **Storage**: append-only record log, LZ4-compressed in 256 KiB blocks
|
||
(XxHash3-checked, `fsync` per write, torn-tail tolerant: a crash
|
||
mid-append truncates cleanly, interior corruption is rejected) with
|
||
in-memory indexes rebuilt on open and automatic compaction (rewrite +
|
||
atomic rename when the log grows past `--compact-threshold`, default
|
||
16 MB). Killed mid-write (`kill -9`), the database recovers all
|
||
committed writes; the log and compaction both work with relative or
|
||
absolute `--db` paths. Records up to the announced 16 MB
|
||
`maxBsonObjectSize` replay correctly.
|
||
- **Concurrency**: a writer-preferring read/write lock splits command
|
||
execution — reads (`find`, `count`, `aggregate`, `list*`) run concurrently
|
||
across connections, writes (CRUD, DDL) are exclusive and totally ordered,
|
||
and handshake/no-op commands run lock-free. The log append + `fsync` still
|
||
happen under the write lock, so the crash guarantees are unchanged. Fine
|
||
for light workloads.
|
||
|
||
## Layout
|
||
|
||
```
|
||
src/
|
||
bson.zig BSON parse/serialize, ObjectId, canonical comparison order
|
||
wire.zig OP_MSG/OP_QUERY framing, message + reply builders
|
||
commands.zig command dispatch (hello, CRUD, aggregate, admin, indexes)
|
||
server.zig TCP accept loop, per-connection handlers, TTL sweep monitor
|
||
db.zig in-memory engine: db → collection → _id → document maps
|
||
storage.zig append-only log: records, replay, CRC validation
|
||
query.zig filter matcher, regex engine, sort, projection
|
||
index.zig secondary indexes: entries, search, query planner
|
||
update.zig update operators with dot-path navigation
|
||
main.zig CLI: --port, --bind, --db, --ttl-sweep-secs, --compact-threshold
|
||
```
|
||
|
||
## Indexes
|
||
|
||
`collection.createIndex({field: 1})` works against every driver; the index
|
||
is persisted in the log, survives restarts and compaction, and is used by
|
||
the query planner to narrow scans.
|
||
|
||
- **Key patterns**: single-field and compound (up to 32 fields), each key
|
||
`1` or `-1`. Descending order is metadata (entries are always stored
|
||
value-ascending); the default index name is MongoDB's `a_1_b_-1`.
|
||
`createIndex({_id: 1})` is an idempotent no-op — the docs map is the
|
||
`_id_` index — and `dropIndex("_id_")` errors.
|
||
- **Options**: `unique` (a conflicting write fails with E11000 naming the
|
||
index; per-document entries are deduped first, so `{a: [1,1]}` is legal)
|
||
and `sparse` (documents missing an indexed field are skipped).
|
||
- **TTL**: `createIndex({expireAt: 1}, {expireAfterSeconds: 60})` deletes a
|
||
document once its indexed date is that many seconds old. A background
|
||
sweeper runs every `--ttl-sweep-secs` seconds (default 60, `0` disables
|
||
it) and deletes through the ordinary write path, so each expiry is logged
|
||
and fsynced and holds across a restart. As in MongoDB the option is
|
||
single-field only (a compound key is `CannotCreateIndex`, code 67),
|
||
`expireAfterSeconds` must be a whole number in `[0, 2147483647]` (`0`
|
||
means "expire at the stored instant"), a non-date value at the path never
|
||
expires, an array of dates expires on its earliest member, and expiry is
|
||
coarse: a document stays visible until the next sweep. Re-creating an
|
||
index with a different expiry is `IndexOptionsConflict` (85) and an
|
||
expiry on `{_id: 1}` is `InvalidIndexSpecificationOption` (197), both as
|
||
MongoDB has them.
|
||
- **Multikey**: an array at an indexed path is indexed as a whole *and*
|
||
element-wise, mirroring the query matcher exactly, so both
|
||
`{tags: "a"}` and `{tags: ["a","b"]}` hit the index. A compound index
|
||
over two array paths rejects the document with MongoDB's "cannot index
|
||
parallel arrays".
|
||
- **Planner**: picks the index covering the longest leading run of
|
||
equality/`$in` predicates (cartesian product capped at 100 lookups),
|
||
optionally with a range on the next key. Ranges with both bounds fall
|
||
back to a scan on multikey indexes (a doc with `{a: [1,2]}` can satisfy
|
||
`{a: {$gt: 5, $lt: 25}}` across two entries), and sparse indexes are
|
||
never used for `null`-valued predicates. The `_id_` fast path resolves
|
||
`{_id: ...}` through the docs map unless the value's compare class is
|
||
serialization-ambiguous (int32 1, int64 1, double 1.0 compare equal but
|
||
hash differently — those fall back to a scan, as do string/symbol/code).
|
||
|
||
v1 limits: no hashed/text/geo/partial indexes, and entry *insert* is O(n)
|
||
(a sorted array memmoves the tail) — fine for a light database, with a
|
||
B-tree as the follow-up. Removal is no longer a scan: entry generation is
|
||
a pure function of the document, so the entries to drop are regenerated
|
||
and found by binary search. A TTL sweep
|
||
walks every entry of every TTL index and holds the write lock for the
|
||
whole pass, so the interval is the tuning knob: expiry is never more
|
||
precise than `--ttl-sweep-secs`, and a very large TTL index wants a
|
||
longer one.
|
||
|
||
## Not (yet) implemented
|
||
|
||
- Authentication (SCRAM) — run without credentials
|
||
- Real cursors (all results are returned in one batch, cursor id 0)
|
||
- Transactions, change streams, replicasets
|
||
- Compression (OP_COMPRESSED)
|
||
- `collMod`, so an index's `expireAfterSeconds` cannot be changed in
|
||
place — drop the index and re-create it with the new expiry
|
||
- `dropCollection`/`dropDatabase` write no log record, so a dropped
|
||
collection (and its index definitions) resurrect on restart
|
||
|
||
## Working with large collections
|
||
|
||
Everything lives in RAM (db → collection → _id → document maps) and every
|
||
write command is logged with `fsync` before it is acknowledged (one sync per
|
||
command via group commit — a 500-doc `insertMany` syncs once, not 500
|
||
times), so multi-GB collections work, with cost/behavior notes measured by
|
||
the `tests/e2e/big.js` harness (12-core/32 GB Mac):
|
||
|
||
- **Build in ReleaseFast** — `zig build` defaults to it. A Debug server is
|
||
10-200x slower on every path (the matcher alone was 70 µs/doc in Debug
|
||
vs 0.4 µs in ReleaseFast), which dwarfed every other difference in the
|
||
MongoDB comparison below.
|
||
- **Compaction no longer needs tuning for bulk loads.** It triggers on the
|
||
share of the log that is garbage rather than on bytes appended, so a pure
|
||
insert workload — which has no garbage — is never rewritten, and a
|
||
rewrite-heavy one is reclaimed once about a fifth of the log is dead,
|
||
keeping the file near 1.25x the live data. `--compact-threshold` is now
|
||
only a floor below which small logs are left alone. (It used to fire
|
||
every 16 MB regardless, rewriting the whole log each time: quadratic
|
||
total traffic, and the reason bulk loads needed a raised threshold.)
|
||
- **`findOne({_id})` is O(1) only for ObjectId ids.** Integer, int64 and
|
||
double ids compare equal but hash differently, so the docs-map fast path
|
||
is skipped and every `_id` lookup becomes a full scan. Use the driver's
|
||
default ObjectIds (or a secondary index) on big collections. The
|
||
order-preserving key encoding already removes the ambiguity that forces
|
||
this; lifting the restriction waits on an ordered `_id` index.
|
||
- **Secondary-index entry insert is O(n)** (sorted array — see v1 limits
|
||
above), so inserting into a collection that already has an index is
|
||
quadratic. Building an index over existing data is not: entries are
|
||
appended unsorted and ordered once. Still cheapest to create indexes
|
||
after the load.
|
||
|
||
## Performance vs MongoDB
|
||
|
||
`tests/e2e/compare-run.sh` runs the same driver workload (1 GB, 65,536 ×
|
||
16 KB docs, every write durable — mongo-lite fsyncs per command, mongod
|
||
runs with `j: true`) against each server and prints a side-by-side table.
|
||
With the ReleaseFast default build (MongoDB 8.3.7 on the same Mac):
|
||
|
||
| benchmark | mongo-lite | mongodb | winner |
|
||
|---|---|---|---|
|
||
| insertOne (sequential) | 0.17 ms | 4.2 ms | **mongo-lite ×25** |
|
||
| bulk insert (insertMany) | 722 MB/s | 824 MB/s | mongodb ×1.1 |
|
||
| createIndex({k: 1}) | 54 ms | 85 ms | **mongo-lite** |
|
||
| countDocuments({}) | 1.5 ms | 11.5 ms | **mongo-lite ×7** |
|
||
| findOne({_id}) | 0.57 ms | 0.50 ms | mongodb |
|
||
| findOne indexed | 0.77 ms | 0.86 ms | mongo-lite |
|
||
| range-scan count | 22.5 ms | 14.5 ms | mongodb ×1.6 |
|
||
| sort + limit(20), on `_id` | 2.5 ms | 2.3 ms | mongodb ×1.1 |
|
||
| sort + limit(20), indexed field | 1.0 ms | — | — |
|
||
| aggregate $group | 10.2 ms | 15.4 ms | **mongo-lite** |
|
||
| updateOne({_id}) | 0.13 ms | 0.22 ms | **mongo-lite** |
|
||
| updateMany (65 docs) | 2.4 ms | 5.8 ms | **mongo-lite ×2.4** |
|
||
| deleteOne + insert | 0.64 ms | 3.9 ms | **mongo-lite ×6** |
|
||
| server RSS | 2.0 GB | 1.6 GB | mongodb (×0.8) |
|
||
| kill -9 → reopen | 0.8 s | 1.3 s | **mongo-lite** |
|
||
| db on disk | 97 MB | 104 MB | **mongo-lite** |
|
||
|
||
The log is now LZ4-compressed in 256 KiB blocks, so the on-disk size is
|
||
on par with MongoDB's compressed files. The remaining losses are
|
||
structural rather than incidental. RSS trails because every document
|
||
carries its own arena and a second full copy as a `Pair` tree; the
|
||
range-scan gap is not the matcher — it is walking 65,536 documents that
|
||
each live in a separate allocation, one pointer chase apiece. And bulk
|
||
insert is compress-bound (the LZ4 codec runs at ~1.7 GB/s; deflate would
|
||
cap writes below the insert rate, which is why the roadmap chose LZ4).
|
||
|
||
Reproduce the table with `bash tests/e2e/compare-run.sh 1g 16k`; the
|
||
pre-tree baseline is recorded in `tests/e2e/results/phase1.txt`, and the
|
||
runs with the B+tree, ordered `_id` index and compressed log (roadmap
|
||
items 1–3) in `tests/e2e/results/phase2.txt`, `phase3.txt` and
|
||
`phase4.txt`.
|
||
|
||
### What is left (highest impact first)
|
||
|
||
Each is written up with its design decisions, ordering constraints and
|
||
traps in [ROADMAP.md](ROADMAP.md).
|
||
|
||
1. **Stop giving every document its own arena** — the source of both the
|
||
RSS gap and the range-scan gap. Storing canonical BSON bytes in a
|
||
per-collection slab and matching against them (parsing only the fields a
|
||
filter names) makes scans contiguous instead of a pointer chase.
|
||
2. **Decompose the global lock** — one reader/writer lock covers the whole
|
||
engine and is held across fsync, compaction and reply construction.
|
||
Per-collection locks plus cross-connection group commit are the path to
|
||
using more than one core on writes.
|
||
|
||
Done so far, with the measurement that drove each:
|
||
|
||
- **Record integrity hash CRC32 → XxHash3.** `std.hash.Crc32` is
|
||
table-driven and byte-at-a-time: 408 MB/s against XxHash3's 31 GB/s, or
|
||
38 µs versus 0.5 µs on a 16 KB document — about two thirds of the entire
|
||
bulk-insert cost. Insert 260 → 700 MB/s.
|
||
- **Compaction triggers on garbage, not on bytes written**, and syncs once
|
||
per rewrite instead of once per document. Bulk load at the default
|
||
threshold 41.6 → 703 MB/s.
|
||
- **Index builds append then sort once** instead of inserting into a sorted
|
||
array. `createIndex` over 65,536 documents 649 → 44 ms.
|
||
- **Index entries hold encoded byte keys**, so comparing them is a memcmp
|
||
rather than a walk over values in unrelated arenas.
|
||
- **A B+tree over the encoded keys** (roadmap item 1): fixed 4 KiB slotted
|
||
pages in a flat u32-addressed node array, an overflow slab for long
|
||
records, no rebalancing on delete, and bulk bottom-up packing. Entry
|
||
insertion and removal are a descent plus a leaf-local edit instead of a
|
||
tail memmove, so writes into an already-built index stopped being
|
||
quadratic. `updateMany` 17.3 → 1.6 ms (2.8x slower than MongoDB → 4x
|
||
faster); `createIndex` 62 → 51 ms.
|
||
- **An ordered `_id` index** (roadmap item 2): every collection carries an
|
||
implicit `_id_` index (kept out of the secondary list, so the listing,
|
||
drop and log-format surfaces are unchanged; rebuilt after replay like
|
||
the secondaries). Its encoded keys are canonical, so the old
|
||
serialization-guarded docs-map fast path is gone and integer/string
|
||
`_id` point lookups, `$in` and ranges hit the tree instead of a full
|
||
scan. `sort({_id: ...})` is now an index-ordered scan with an early stop:
|
||
`sort+limit(20)` 6.2 → 2.4 ms (parity with MongoDB).
|
||
- **A block-framed, LZ4-compressed log** (roadmap item 3): a file header
|
||
plus ~256 KiB blocks, each holding the existing record framing with the
|
||
integrity hash covering the stored bytes (so the decompressor only ever
|
||
sees input already proven intact). Records never straddle blocks; a
|
||
short read, impossible length or hash mismatch in the final block is a
|
||
torn tail (truncate cleanly), anywhere else is corruption. The hand-rolled
|
||
LZ4 codec runs at ~1.7 GB/s and falls back to raw per block when
|
||
compression does not help. `db on disk` 1025 → 97 MB — now smaller than
|
||
MongoDB's own compressed files.
|
||
- **Entry removal is a binary search**, not a scan of the whole index.
|
||
`updateMany` 15.4 → 5.5 ms.
|
||
- **Top-k sort selection** and an allocation-free decorate pass, plus
|
||
**index-supplied ordering** when an index already holds candidates in the
|
||
requested order. `sort+limit(20)` 40 → 4.3 ms, or 1.0 ms on an indexed
|
||
field.
|
||
- **`limit` reaches the scan**, which used to materialize the whole
|
||
collection before slicing, and `countDocuments` is answered by counting
|
||
rather than by materializing and discarding every match.
|
||
- **Matching collects candidates on the stack**, resolves operators to an
|
||
enum once per filter field rather than by string per document, and reuses
|
||
one reply arena per connection.
|
||
|
||
Several real bugs surfaced while benchmarking:
|
||
|
||
- `plan_id` returned a pointer to a stack temporary (`&.{e}`) that dangled
|
||
after the frame returned — Debug tolerated it, ReleaseFast read garbage,
|
||
silently breaking every `findOne({_id: <ObjectId>})`. It now heap-copies
|
||
the lookup value and frees it.
|
||
- Multi-doc writes fsynced once per document; they now group-commit (one
|
||
fsync per command, same crash guarantees — verified by the kill -9
|
||
crash suites).
|
||
- Compaction fsynced once per live document, because the log it wrote into
|
||
never had deferred syncing enabled — 65,536 fsyncs to rewrite a 1 GB
|
||
collection.
|
||
- `remove` never checked the compaction threshold, so a delete-heavy
|
||
workload grew the log without bound.
|
||
|
||
|
||
## Code style
|
||
|
||
Zig 0.16 idioms (`std.Io` threaded through everything, unmanaged
|
||
containers); user-declared functions use `snake_case` per this repo's house
|
||
style.
|