Five items in dependency order, each sized to land on its own. The design decisions already settled are recorded so they are not re-derived, and so are the ordering constraints, which are the part that actually matters -- notably that the _id index must follow the tree, because it updates on every insert and against a sorted array that is only affordable while ids happen to append at the end. Also carries the ground rules the earlier work established: A/B on one harness rather than trusting a model, mutation-check tests that guard an invariant, and the two absolutes (an index may only over-approximate; the database must always open). Each item names the traps found while investigating it -- deletion being where B+trees go wrong, listIndexes possibly noticing a real _id_ index, hashing compressed rather than uncompressed bytes, the fabricated Document values that break when Document changes meaning, and the durability guarantee that quietly weakens under group commit.
mongo-lite
A lightweight, embedded MongoDB-compatible document database written in
Zig 0.16. Like SQLite, it stores everything in a single file; unlike SQLite,
it speaks the MongoDB wire protocol, so real clients — mongosh, the Node.js
driver, PyMongo — connect over TCP and just work.
Quick start
zig build # build the server
zig build test # run the unit test suite
zig-out/bin/mongo-lite --port 27017 --db data.log --compact-threshold 256m
# in another terminal:
mongosh --port 27017
> db.users.insertOne({name: "alice", age: 30})
> db.users.find({age: {$gt: 25}}).toArray()
> db.users.updateOne({name: "alice"}, {$set: {vip: true}})
> db.users.deleteOne({name: "bob"})
> db.sessions.createIndex({expireAt: 1}, {expireAfterSeconds: 3600})
Features
- Wire protocol: OP_MSG (2013) plus legacy OP_QUERY/OP_REPLY (2004/2001)
for the driver handshake; hello/isMaster with
maxWireVersion: 8, so modern drivers (Node, Python, mongosh) connect without workarounds. - BSON: full parse/serialize round-trip for all common types (including binary, regex, timestamps, ObjectId), canonical MongoDB comparison order for sorting and range queries.
- CRUD:
insert,find(filter, sort, skip/limit, projection),update(multi/upsert),delete,findAndModify,count,aggregate($match,$sort,$skip,$limit,$project,$count,$groupwith$sum), pluscreate/drop/listCollections/listDatabases/dropDatabase. - Query operators:
$eq$ne$gt$gte$lt$lte$in$nin$exists$regex(hand-rolled engine: anchors,.,* + ?, character classes, groups, alternation,i/soptions)$not$and$or$nor$size$all$elemMatch, with dot paths and array multikey semantics. - Secondary indexes:
createIndex/listIndexes/dropIndexvia the three driver commands, single-field and compound, withunique,sparseandexpireAfterSeconds(TTL) options, persisted in the log and rebuilt on open (compaction re-emits them). A background sweeper expires TTL-indexed documents through the ordinary logged write path. The query planner turns equality /$in/ range predicates into index lookups acrossfind,count,update,delete,findAndModify, and a leading$matchinaggregate; every candidate is re-checked against the full filter, so an index that over-approximates is merely slow, never wrong. - Update operators:
$set$unset$inc$push($each)$pull$rename, with dot-path creation (including array indices). - Storage: append-only record log (CRC32-checked,
fsyncper write, torn-tail tolerant) with in-memory indexes rebuilt on open and automatic compaction (rewrite + atomic rename when the log grows past--compact-threshold, default 16 MB). Killed mid-write (kill -9), the database recovers all committed writes; the log and compaction both work with relative or absolute--dbpaths. Records up to the announced 16 MBmaxBsonObjectSizereplay correctly. - Concurrency: a writer-preferring read/write lock splits command
execution — reads (
find,count,aggregate,list*) run concurrently across connections, writes (CRUD, DDL) are exclusive and totally ordered, and handshake/no-op commands run lock-free. The log append +fsyncstill happen under the write lock, so the crash guarantees are unchanged. Fine for light workloads.
Layout
src/
bson.zig BSON parse/serialize, ObjectId, canonical comparison order
wire.zig OP_MSG/OP_QUERY framing, message + reply builders
commands.zig command dispatch (hello, CRUD, aggregate, admin, indexes)
server.zig TCP accept loop, per-connection handlers, TTL sweep monitor
db.zig in-memory engine: db → collection → _id → document maps
storage.zig append-only log: records, replay, CRC validation
query.zig filter matcher, regex engine, sort, projection
index.zig secondary indexes: entries, search, query planner
update.zig update operators with dot-path navigation
main.zig CLI: --port, --bind, --db, --ttl-sweep-secs, --compact-threshold
Indexes
collection.createIndex({field: 1}) works against every driver; the index
is persisted in the log, survives restarts and compaction, and is used by
the query planner to narrow scans.
- Key patterns: single-field and compound (up to 32 fields), each key
1or-1. Descending order is metadata (entries are always stored value-ascending); the default index name is MongoDB'sa_1_b_-1.createIndex({_id: 1})is an idempotent no-op — the docs map is the_id_index — anddropIndex("_id_")errors. - Options:
unique(a conflicting write fails with E11000 naming the index; per-document entries are deduped first, so{a: [1,1]}is legal) andsparse(documents missing an indexed field are skipped). - TTL:
createIndex({expireAt: 1}, {expireAfterSeconds: 60})deletes a document once its indexed date is that many seconds old. A background sweeper runs every--ttl-sweep-secsseconds (default 60,0disables it) and deletes through the ordinary write path, so each expiry is logged and fsynced and holds across a restart. As in MongoDB the option is single-field only (a compound key isCannotCreateIndex, code 67),expireAfterSecondsmust be a whole number in[0, 2147483647](0means "expire at the stored instant"), a non-date value at the path never expires, an array of dates expires on its earliest member, and expiry is coarse: a document stays visible until the next sweep. Re-creating an index with a different expiry isIndexOptionsConflict(85) and an expiry on{_id: 1}isInvalidIndexSpecificationOption(197), both as MongoDB has them. - Multikey: an array at an indexed path is indexed as a whole and
element-wise, mirroring the query matcher exactly, so both
{tags: "a"}and{tags: ["a","b"]}hit the index. A compound index over two array paths rejects the document with MongoDB's "cannot index parallel arrays". - Planner: picks the index covering the longest leading run of
equality/
$inpredicates (cartesian product capped at 100 lookups), optionally with a range on the next key. Ranges with both bounds fall back to a scan on multikey indexes (a doc with{a: [1,2]}can satisfy{a: {$gt: 5, $lt: 25}}across two entries), and sparse indexes are never used fornull-valued predicates. The_id_fast path resolves{_id: ...}through the docs map unless the value's compare class is serialization-ambiguous (int32 1, int64 1, double 1.0 compare equal but hash differently — those fall back to a scan, as do string/symbol/code).
v1 limits: no hashed/text/geo/partial indexes, and entry insert is O(n)
(a sorted array memmoves the tail) — fine for a light database, with a
B-tree as the follow-up. Removal is no longer a scan: entry generation is
a pure function of the document, so the entries to drop are regenerated
and found by binary search. A TTL sweep
walks every entry of every TTL index and holds the write lock for the
whole pass, so the interval is the tuning knob: expiry is never more
precise than --ttl-sweep-secs, and a very large TTL index wants a
longer one.
Not (yet) implemented
- Authentication (SCRAM) — run without credentials
- Real cursors (all results are returned in one batch, cursor id 0)
- Transactions, change streams, replicasets
- Compression (OP_COMPRESSED)
collMod, so an index'sexpireAfterSecondscannot be changed in place — drop the index and re-create it with the new expirydropCollection/dropDatabasewrite no log record, so a dropped collection (and its index definitions) resurrect on restart
Working with large collections
Everything lives in RAM (db → collection → _id → document maps) and every
write command is logged with fsync before it is acknowledged (one sync per
command via group commit — a 500-doc insertMany syncs once, not 500
times), so multi-GB collections work, with cost/behavior notes measured by
the tests/e2e/big.js harness (12-core/32 GB Mac):
- Build in ReleaseFast —
zig builddefaults to it. A Debug server is 10-200x slower on every path (the matcher alone was 70 µs/doc in Debug vs 0.4 µs in ReleaseFast), which dwarfed every other difference in the MongoDB comparison below. - Compaction no longer needs tuning for bulk loads. It triggers on the
share of the log that is garbage rather than on bytes appended, so a pure
insert workload — which has no garbage — is never rewritten, and a
rewrite-heavy one is reclaimed once about a fifth of the log is dead,
keeping the file near 1.25x the live data.
--compact-thresholdis now only a floor below which small logs are left alone. (It used to fire every 16 MB regardless, rewriting the whole log each time: quadratic total traffic, and the reason bulk loads needed a raised threshold.) findOne({_id})is O(1) only for ObjectId ids. Integer, int64 and double ids compare equal but hash differently, so the docs-map fast path is skipped and every_idlookup becomes a full scan. Use the driver's default ObjectIds (or a secondary index) on big collections. The order-preserving key encoding already removes the ambiguity that forces this; lifting the restriction waits on an ordered_idindex.- Secondary-index entry insert is O(n) (sorted array — see v1 limits above), so inserting into a collection that already has an index is quadratic. Building an index over existing data is not: entries are appended unsorted and ordered once. Still cheapest to create indexes after the load.
Performance vs MongoDB
tests/e2e/compare-run.sh runs the same driver workload (1 GB, 65,536 ×
16 KB docs, every write durable — mongo-lite fsyncs per command, mongod
runs with j: true) against each server and prints a side-by-side table.
With the ReleaseFast default build (MongoDB 8.3.7 on the same Mac):
| benchmark | mongo-lite | mongodb | winner |
|---|---|---|---|
| insertOne (sequential) | 0.19 ms | 5.0 ms | mongo-lite ×26 |
| bulk insert (insertMany) | 853 MB/s | 690 MB/s | mongo-lite ×1.2 |
| createIndex({k: 1}) | 62 ms | 78 ms | mongo-lite |
| countDocuments({}) | 1.5 ms | 11.5 ms | mongo-lite ×8 |
| findOne({_id}) | 0.48 ms | 0.54 ms | mongo-lite |
| findOne indexed | 0.64 ms | 4.3 ms | mongo-lite ×7 |
| range-scan count | 21 ms | 13 ms | mongodb ×1.7 |
sort + limit(20), on _id |
4.3 ms | 2.0 ms | mongodb ×2 |
| sort + limit(20), indexed field | 1.0 ms | — | — |
| aggregate $group | 9.8 ms | 13.7 ms | mongo-lite |
| updateOne({_id}) | 0.16 ms | 0.19 ms | mongo-lite |
| updateMany (65 docs) | 5.5 ms | 6.1 ms | mongo-lite |
| deleteOne + insert | 0.62 ms | 4.9 ms | mongo-lite ×8 |
| server RSS | 2.0 GB | 1.4 GB | mongodb (×0.7) |
| kill -9 → reopen | 0.8 s | 1.3 s | mongo-lite |
| db on disk | 1.0 GB | 93 MB | mongodb (compressed) |
The remaining losses are structural rather than incidental. Disk size is
the big one: payloads are stored raw, so the log is 11x MongoDB's
compressed files. RSS trails because every document carries its own arena.
The range-scan gap is not the matcher — it is walking 65,536 documents
that each live in a separate allocation, one pointer chase apiece. And
sort on _id still materializes candidates because nothing ordered
covers _id yet; the same sort on an indexed field streams straight out
of the index at 1.0 ms.
Reproduce the table with bash tests/e2e/compare-run.sh 1g 16k; the run
above is recorded in tests/e2e/results/phase1.txt.
What is left (highest impact first)
Each is written up with its design decisions, ordering constraints and traps in ROADMAP.md.
- Compress the log — the largest remaining gap (×11). Payloads are stored raw. A block-framed format with an LZ4 block codec would shrink highly compressible workloads massively; note Zig 0.16 ships zstd decompression only, and deflate would cap writes below the current insert rate.
- A B-tree over the encoded keys — entry insert still memmoves the
tail of a sorted array, so writing into a collection that already has an
index is quadratic. A flat,
u32-indexed node array would also be dumpable into a checkpoint, which is what makes a fast reopen possible. - An ordered
_idindex —sort({_id: ...})still materializes every candidate, and integer_ids still scan. Both fall out of indexing the encoded_id. It wants the tree first: an_idindex updates on every insert, and doing that against a sorted array is only cheap because ObjectIds append at the end. - Stop giving every document its own arena — the source of both the RSS gap and the range-scan gap. Storing canonical BSON bytes in a per-collection slab and matching against them (parsing only the fields a filter names) makes scans contiguous instead of a pointer chase.
- Decompose the global lock — one reader/writer lock covers the whole engine and is held across fsync, compaction and reply construction. Per-collection locks plus cross-connection group commit are the path to using more than one core on writes.
Done so far, with the measurement that drove each:
- Record integrity hash CRC32 → XxHash3.
std.hash.Crc32is table-driven and byte-at-a-time: 408 MB/s against XxHash3's 31 GB/s, or 38 µs versus 0.5 µs on a 16 KB document — about two thirds of the entire bulk-insert cost. Insert 260 → 700 MB/s. - Compaction triggers on garbage, not on bytes written, and syncs once per rewrite instead of once per document. Bulk load at the default threshold 41.6 → 703 MB/s.
- Index builds append then sort once instead of inserting into a sorted
array.
createIndexover 65,536 documents 649 → 44 ms. - Index entries hold encoded byte keys, so comparing them is a memcmp rather than a walk over values in unrelated arenas.
- Entry removal is a binary search, not a scan of the whole index.
updateMany15.4 → 5.5 ms. - Top-k sort selection and an allocation-free decorate pass, plus
index-supplied ordering when an index already holds candidates in the
requested order.
sort+limit(20)40 → 4.3 ms, or 1.0 ms on an indexed field. limitreaches the scan, which used to materialize the whole collection before slicing, andcountDocumentsis answered by counting rather than by materializing and discarding every match.- Matching collects candidates on the stack, resolves operators to an enum once per filter field rather than by string per document, and reuses one reply arena per connection.
Several real bugs surfaced while benchmarking:
plan_idreturned a pointer to a stack temporary (&.{e}) that dangled after the frame returned — Debug tolerated it, ReleaseFast read garbage, silently breaking everyfindOne({_id: <ObjectId>}). It now heap-copies the lookup value and frees it.- Multi-doc writes fsynced once per document; they now group-commit (one fsync per command, same crash guarantees — verified by the kill -9 crash suites).
- Compaction fsynced once per live document, because the log it wrote into never had deferred syncing enabled — 65,536 fsyncs to rewrite a 1 GB collection.
removenever checked the compaction threshold, so a delete-heavy workload grew the log without bound.
Code style
Zig 0.16 idioms (std.Io threaded through everything, unmanaged
containers); user-declared functions use snake_case per this repo's house
style.