# mongo-lite A lightweight, embedded MongoDB-compatible document database written in Zig 0.16. Like SQLite, it stores everything in a single file; unlike SQLite, it speaks the MongoDB wire protocol, so real clients — `mongosh`, the Node.js driver, PyMongo — connect over TCP and just work. ## Quick start ```sh zig build # build the server zig build test # run the unit test suite zig-out/bin/mongo-lite --port 27017 --db data.log --compact-threshold 256m # in another terminal: mongosh --port 27017 > db.users.insertOne({name: "alice", age: 30}) > db.users.find({age: {$gt: 25}}).toArray() > db.users.updateOne({name: "alice"}, {$set: {vip: true}}) > db.users.deleteOne({name: "bob"}) > db.sessions.createIndex({expireAt: 1}, {expireAfterSeconds: 3600}) ``` ## Features - **Wire protocol**: OP_MSG (2013) plus legacy OP_QUERY/OP_REPLY (2004/2001) for the driver handshake; hello/isMaster with `maxWireVersion: 8`, so modern drivers (Node, Python, mongosh) connect without workarounds. - **BSON**: full parse/serialize round-trip for all common types (including binary, regex, timestamps, ObjectId), canonical MongoDB comparison order for sorting and range queries. - **CRUD**: `insert`, `find` (filter, sort, skip/limit, projection), `update` (multi/upsert), `delete`, `findAndModify`, `count`, `aggregate` (`$match`, `$sort`, `$skip`, `$limit`, `$project`, `$count`, `$group` with `$sum`), plus `create`/`drop`/`listCollections`/ `listDatabases`/`dropDatabase`. - **Query operators**: `$eq` `$ne` `$gt` `$gte` `$lt` `$lte` `$in` `$nin` `$exists` `$regex` (hand-rolled engine: anchors, `.`, `* + ?`, character classes, groups, alternation, `i`/`s` options) `$not` `$and` `$or` `$nor` `$size` `$all` `$elemMatch`, with dot paths and array multikey semantics. - **Secondary indexes**: `createIndex`/`listIndexes`/`dropIndex` via the three driver commands, single-field and compound, with `unique`, `sparse` and `expireAfterSeconds` (TTL) options, persisted in the log and rebuilt on open (compaction re-emits them). A background sweeper expires TTL-indexed documents through the ordinary logged write path. The query planner turns equality / `$in` / range predicates into index lookups across `find`, `count`, `update`, `delete`, `findAndModify`, and a leading `$match` in `aggregate`; every candidate is re-checked against the full filter, so an index that over-approximates is merely slow, never wrong. - **Update operators**: `$set` `$unset` `$inc` `$push` (`$each`) `$pull` `$rename`, with dot-path creation (including array indices). - **Storage**: append-only record log (CRC32-checked, `fsync` per write, torn-tail tolerant) with in-memory indexes rebuilt on open and automatic compaction (rewrite + atomic rename when the log grows past `--compact-threshold`, default 16 MB). Killed mid-write (`kill -9`), the database recovers all committed writes; the log and compaction both work with relative or absolute `--db` paths. Records up to the announced 16 MB `maxBsonObjectSize` replay correctly. - **Concurrency**: a writer-preferring read/write lock splits command execution — reads (`find`, `count`, `aggregate`, `list*`) run concurrently across connections, writes (CRUD, DDL) are exclusive and totally ordered, and handshake/no-op commands run lock-free. The log append + `fsync` still happen under the write lock, so the crash guarantees are unchanged. Fine for light workloads. ## Layout ``` src/ bson.zig BSON parse/serialize, ObjectId, canonical comparison order wire.zig OP_MSG/OP_QUERY framing, message + reply builders commands.zig command dispatch (hello, CRUD, aggregate, admin, indexes) server.zig TCP accept loop, per-connection handlers, TTL sweep monitor db.zig in-memory engine: db → collection → _id → document maps storage.zig append-only log: records, replay, CRC validation query.zig filter matcher, regex engine, sort, projection index.zig secondary indexes: entries, search, query planner update.zig update operators with dot-path navigation main.zig CLI: --port, --bind, --db, --ttl-sweep-secs, --compact-threshold ``` ## Indexes `collection.createIndex({field: 1})` works against every driver; the index is persisted in the log, survives restarts and compaction, and is used by the query planner to narrow scans. - **Key patterns**: single-field and compound (up to 32 fields), each key `1` or `-1`. Descending order is metadata (entries are always stored value-ascending); the default index name is MongoDB's `a_1_b_-1`. `createIndex({_id: 1})` is an idempotent no-op — the docs map is the `_id_` index — and `dropIndex("_id_")` errors. - **Options**: `unique` (a conflicting write fails with E11000 naming the index; per-document entries are deduped first, so `{a: [1,1]}` is legal) and `sparse` (documents missing an indexed field are skipped). - **TTL**: `createIndex({expireAt: 1}, {expireAfterSeconds: 60})` deletes a document once its indexed date is that many seconds old. A background sweeper runs every `--ttl-sweep-secs` seconds (default 60, `0` disables it) and deletes through the ordinary write path, so each expiry is logged and fsynced and holds across a restart. As in MongoDB the option is single-field only (a compound key is `CannotCreateIndex`, code 67), `expireAfterSeconds` must be a whole number in `[0, 2147483647]` (`0` means "expire at the stored instant"), a non-date value at the path never expires, an array of dates expires on its earliest member, and expiry is coarse: a document stays visible until the next sweep. Re-creating an index with a different expiry is `IndexOptionsConflict` (85) and an expiry on `{_id: 1}` is `InvalidIndexSpecificationOption` (197), both as MongoDB has them. - **Multikey**: an array at an indexed path is indexed as a whole *and* element-wise, mirroring the query matcher exactly, so both `{tags: "a"}` and `{tags: ["a","b"]}` hit the index. A compound index over two array paths rejects the document with MongoDB's "cannot index parallel arrays". - **Planner**: picks the index covering the longest leading run of equality/`$in` predicates (cartesian product capped at 100 lookups), optionally with a range on the next key. Ranges with both bounds fall back to a scan on multikey indexes (a doc with `{a: [1,2]}` can satisfy `{a: {$gt: 5, $lt: 25}}` across two entries), and sparse indexes are never used for `null`-valued predicates. The `_id_` fast path resolves `{_id: ...}` through the docs map unless the value's compare class is serialization-ambiguous (int32 1, int64 1, double 1.0 compare equal but hash differently — those fall back to a scan, as do string/symbol/code). v1 limits: no hashed/text/geo/partial indexes, and entry *insert* is O(n) (a sorted array memmoves the tail) — fine for a light database, with a B-tree as the follow-up. Removal is no longer a scan: entry generation is a pure function of the document, so the entries to drop are regenerated and found by binary search. A TTL sweep walks every entry of every TTL index and holds the write lock for the whole pass, so the interval is the tuning knob: expiry is never more precise than `--ttl-sweep-secs`, and a very large TTL index wants a longer one. ## Not (yet) implemented - Authentication (SCRAM) — run without credentials - Real cursors (all results are returned in one batch, cursor id 0) - Transactions, change streams, replicasets - Compression (OP_COMPRESSED) - `collMod`, so an index's `expireAfterSeconds` cannot be changed in place — drop the index and re-create it with the new expiry - `dropCollection`/`dropDatabase` write no log record, so a dropped collection (and its index definitions) resurrect on restart ## Working with large collections Everything lives in RAM (db → collection → _id → document maps) and every write command is logged with `fsync` before it is acknowledged (one sync per command via group commit — a 500-doc `insertMany` syncs once, not 500 times), so multi-GB collections work, with cost/behavior notes measured by the `tests/e2e/big.js` harness (12-core/32 GB Mac): - **Build in ReleaseFast** — `zig build` defaults to it. A Debug server is 10-200x slower on every path (the matcher alone was 70 µs/doc in Debug vs 0.4 µs in ReleaseFast), which dwarfed every other difference in the MongoDB comparison below. - **Compaction no longer needs tuning for bulk loads.** It triggers on the share of the log that is garbage rather than on bytes appended, so a pure insert workload — which has no garbage — is never rewritten, and a rewrite-heavy one is reclaimed once about a fifth of the log is dead, keeping the file near 1.25x the live data. `--compact-threshold` is now only a floor below which small logs are left alone. (It used to fire every 16 MB regardless, rewriting the whole log each time: quadratic total traffic, and the reason bulk loads needed a raised threshold.) - **`findOne({_id})` is O(1) only for ObjectId ids.** Integer, int64 and double ids compare equal but hash differently, so the docs-map fast path is skipped and every `_id` lookup becomes a full scan. Use the driver's default ObjectIds (or a secondary index) on big collections. The order-preserving key encoding already removes the ambiguity that forces this; lifting the restriction waits on an ordered `_id` index. - **Secondary-index entry insert is O(n)** (sorted array — see v1 limits above), so inserting into a collection that already has an index is quadratic. Building an index over existing data is not: entries are appended unsorted and ordered once. Still cheapest to create indexes after the load. ## Performance vs MongoDB `tests/e2e/compare-run.sh` runs the same driver workload (1 GB, 65,536 × 16 KB docs, every write durable — mongo-lite fsyncs per command, mongod runs with `j: true`) against each server and prints a side-by-side table. With the ReleaseFast default build (MongoDB 8.3.7 on the same Mac): | benchmark | mongo-lite | mongodb | winner | |---|---|---|---| | insertOne (sequential) | 0.19 ms | 4.1 ms | **mongo-lite ×22** | | bulk insert (insertMany) | 810 MB/s | 714 MB/s | **mongo-lite ×1.1** | | createIndex({k: 1}) | 51 ms | 82 ms | **mongo-lite** | | countDocuments({}) | 1.5 ms | 13.8 ms | **mongo-lite ×9** | | findOne({_id}) | 0.57 ms | 0.67 ms | mongo-lite | | findOne indexed | 0.58 ms | 1.5 ms | **mongo-lite ×2.6** | | range-scan count | 22 ms | 13 ms | mongodb ×1.7 | | sort + limit(20), on `_id` | 2.4 ms | 2.2 ms | mongodb ×1.1 | | sort + limit(20), indexed field | 1.0 ms | — | — | | aggregate $group | 11.5 ms | 15.5 ms | **mongo-lite** | | updateOne({_id}) | 0.17 ms | 0.19 ms | mongo-lite | | updateMany (65 docs) | 1.6 ms | 6.4 ms | **mongo-lite ×4** | | deleteOne + insert | 0.62 ms | 4.8 ms | **mongo-lite ×8** | | server RSS | 2.0 GB | 1.5 GB | mongodb (×0.7) | | kill -9 → reopen | 0.8 s | 1.3 s | **mongo-lite** | | db on disk | 1.0 GB | 96 MB | mongodb (compressed) | The remaining losses are structural rather than incidental. Disk size is the big one: payloads are stored raw, so the log is 11x MongoDB's compressed files. RSS trails because every document carries its own arena. The range-scan gap is not the matcher — it is walking 65,536 documents that each live in a separate allocation, one pointer chase apiece. Reproduce the table with `bash tests/e2e/compare-run.sh 1g 16k`; the pre-tree baseline is recorded in `tests/e2e/results/phase1.txt`, and the runs with the B+tree and ordered `_id` index (roadmap items 1 and 2) in `tests/e2e/results/phase2.txt` and `tests/e2e/results/phase3.txt`. ### What is left (highest impact first) Each is written up with its design decisions, ordering constraints and traps in [ROADMAP.md](ROADMAP.md). 1. **Compress the log** — the largest remaining gap (×11). Payloads are stored raw. A block-framed format with an LZ4 block codec would shrink highly compressible workloads massively; note Zig 0.16 ships zstd decompression only, and deflate would cap writes below the current insert rate. 2. **Stop giving every document its own arena** — the source of both the RSS gap and the range-scan gap. Storing canonical BSON bytes in a per-collection slab and matching against them (parsing only the fields a filter names) makes scans contiguous instead of a pointer chase. 3. **Decompose the global lock** — one reader/writer lock covers the whole engine and is held across fsync, compaction and reply construction. Per-collection locks plus cross-connection group commit are the path to using more than one core on writes. Done so far, with the measurement that drove each: - **Record integrity hash CRC32 → XxHash3.** `std.hash.Crc32` is table-driven and byte-at-a-time: 408 MB/s against XxHash3's 31 GB/s, or 38 µs versus 0.5 µs on a 16 KB document — about two thirds of the entire bulk-insert cost. Insert 260 → 700 MB/s. - **Compaction triggers on garbage, not on bytes written**, and syncs once per rewrite instead of once per document. Bulk load at the default threshold 41.6 → 703 MB/s. - **Index builds append then sort once** instead of inserting into a sorted array. `createIndex` over 65,536 documents 649 → 44 ms. - **Index entries hold encoded byte keys**, so comparing them is a memcmp rather than a walk over values in unrelated arenas. - **A B+tree over the encoded keys** (roadmap item 1): fixed 4 KiB slotted pages in a flat u32-addressed node array, an overflow slab for long records, no rebalancing on delete, and bulk bottom-up packing. Entry insertion and removal are a descent plus a leaf-local edit instead of a tail memmove, so writes into an already-built index stopped being quadratic. `updateMany` 17.3 → 1.6 ms (2.8x slower than MongoDB → 4x faster); `createIndex` 62 → 51 ms. - **An ordered `_id` index** (roadmap item 2): every collection carries an implicit `_id_` index (kept out of the secondary list, so the listing, drop and log-format surfaces are unchanged; rebuilt after replay like the secondaries). Its encoded keys are canonical, so the old serialization-guarded docs-map fast path is gone and integer/string `_id` point lookups, `$in` and ranges hit the tree instead of a full scan. `sort({_id: ...})` is now an index-ordered scan with an early stop: `sort+limit(20)` 6.2 → 2.4 ms (parity with MongoDB). - **Entry removal is a binary search**, not a scan of the whole index. `updateMany` 15.4 → 5.5 ms. - **Top-k sort selection** and an allocation-free decorate pass, plus **index-supplied ordering** when an index already holds candidates in the requested order. `sort+limit(20)` 40 → 4.3 ms, or 1.0 ms on an indexed field. - **`limit` reaches the scan**, which used to materialize the whole collection before slicing, and `countDocuments` is answered by counting rather than by materializing and discarding every match. - **Matching collects candidates on the stack**, resolves operators to an enum once per filter field rather than by string per document, and reuses one reply arena per connection. Several real bugs surfaced while benchmarking: - `plan_id` returned a pointer to a stack temporary (`&.{e}`) that dangled after the frame returned — Debug tolerated it, ReleaseFast read garbage, silently breaking every `findOne({_id: })`. It now heap-copies the lookup value and frees it. - Multi-doc writes fsynced once per document; they now group-commit (one fsync per command, same crash guarantees — verified by the kill -9 crash suites). - Compaction fsynced once per live document, because the log it wrote into never had deferred syncing enabled — 65,536 fsyncs to rewrite a 1 GB collection. - `remove` never checked the compaction threshold, so a delete-heavy workload grew the log without bound. ## Code style Zig 0.16 idioms (`std.Io` threaded through everything, unmanaged containers); user-declared functions use `snake_case` per this repo's house style.