Two streams of work land together: they are interleaved in storage.zig
and db.zig and only build as a unit.
Already in the working tree before this session:
- ReleaseFast as the default zig build (Debug was 10-200x slower)
- group commit: one fsync per write command instead of per document
- plan_id returned a pointer to a stack temporary; ReleaseFast read
garbage and silently broke findOne({_id: ObjectId})
- perf suite: big.js, compare.js, compare-run.sh, e2e6.js
Phase 1 performance work:
Record integrity hash CRC32 -> XxHash3. std.hash.Crc32 is the
table-driven byte-at-a-time Crc32IsoHdlc, measured at 408 MB/s against
XxHash3's 31 GB/s: 38us versus 0.5us on a 16 KiB document, which was
about two thirds of the entire bulk-insert cost. The record header
grows from u32 crc to u64 hash (header_len 20 -> 24), a breaking
format change. Bulk insert 260 -> 700 MB/s, reopen 1.1 -> 0.5s.
Compaction fsynced once per live document, because Log.open leaves
defer_sync false and compact never set it. It now issues one sync for
the whole rewrite, before the rename that publishes it.
Compaction triggers on the share of the log that is garbage
(live_docs/dead_docs, maintained at evict_doc, the single point where
a document dies) rather than on bytes appended. A fixed byte count is
wrong in both directions: a 1 GB bulk load holds no garbage at all yet
would compact ~64 times under the 16 MiB default, rewriting 1 GB each
time, while a small collection rewritten in place accumulates garbage
indefinitely without ever reaching the count. Pure inserts now never
compact, and the file stays near 1.25x the live data. Bulk load at the
default threshold: 41.6 -> 702.7 MB/s.
remove() never called maybe_compact, so a delete-heavy workload grew
the log without bound.
e2e6's compaction check required the file to bloat past 30 MiB before
being reclaimed, which encoded the old policy and failed on strictly
better behaviour (ends at 15.8 MB against ~12 MB live, was ~30 MB). It
now asserts the file ends near the live size and peaked well above it,
which does not depend on when the trigger fires. Sampling interval 50
-> 10ms: the operations now finish inside the old window.
Verified: 70 unit tests under both ReleaseFast and ReleaseSafe, e2e
29/16/17/3/2 checks, the crash-a/kill -9/crash-b pair, and e2e6 72/72
across three consecutive runs.
243 lines
13 KiB
Markdown
243 lines
13 KiB
Markdown
# mongo-lite
|
||
|
||
A lightweight, embedded MongoDB-compatible document database written in
|
||
Zig 0.16. Like SQLite, it stores everything in a single file; unlike SQLite,
|
||
it speaks the MongoDB wire protocol, so real clients — `mongosh`, the Node.js
|
||
driver, PyMongo — connect over TCP and just work.
|
||
|
||
## Quick start
|
||
|
||
```sh
|
||
zig build # build the server
|
||
zig build test # run the unit test suite
|
||
|
||
zig-out/bin/mongo-lite --port 27017 --db data.log --compact-threshold 256m
|
||
|
||
# in another terminal:
|
||
mongosh --port 27017
|
||
> db.users.insertOne({name: "alice", age: 30})
|
||
> db.users.find({age: {$gt: 25}}).toArray()
|
||
> db.users.updateOne({name: "alice"}, {$set: {vip: true}})
|
||
> db.users.deleteOne({name: "bob"})
|
||
> db.sessions.createIndex({expireAt: 1}, {expireAfterSeconds: 3600})
|
||
```
|
||
|
||
## Features
|
||
|
||
- **Wire protocol**: OP_MSG (2013) plus legacy OP_QUERY/OP_REPLY (2004/2001)
|
||
for the driver handshake; hello/isMaster with `maxWireVersion: 8`, so
|
||
modern drivers (Node, Python, mongosh) connect without workarounds.
|
||
- **BSON**: full parse/serialize round-trip for all common types
|
||
(including binary, regex, timestamps, ObjectId), canonical MongoDB
|
||
comparison order for sorting and range queries.
|
||
- **CRUD**: `insert`, `find` (filter, sort, skip/limit, projection),
|
||
`update` (multi/upsert), `delete`, `findAndModify`, `count`,
|
||
`aggregate` (`$match`, `$sort`, `$skip`, `$limit`, `$project`, `$count`,
|
||
`$group` with `$sum`), plus `create`/`drop`/`listCollections`/
|
||
`listDatabases`/`dropDatabase`.
|
||
- **Query operators**: `$eq` `$ne` `$gt` `$gte` `$lt` `$lte` `$in` `$nin`
|
||
`$exists` `$regex` (hand-rolled engine: anchors, `.`, `* + ?`, character
|
||
classes, groups, alternation, `i`/`s` options) `$not` `$and` `$or` `$nor`
|
||
`$size` `$all` `$elemMatch`, with dot paths and array multikey semantics.
|
||
- **Secondary indexes**: `createIndex`/`listIndexes`/`dropIndex` via the
|
||
three driver commands, single-field and compound, with `unique`,
|
||
`sparse` and `expireAfterSeconds` (TTL) options, persisted in the log
|
||
and rebuilt on open (compaction re-emits them). A background sweeper
|
||
expires TTL-indexed documents through the ordinary logged write path.
|
||
The query planner turns equality / `$in` / range
|
||
predicates into index lookups across `find`, `count`, `update`,
|
||
`delete`, `findAndModify`, and a leading `$match` in `aggregate`; every
|
||
candidate is re-checked against the full filter, so an index that
|
||
over-approximates is merely slow, never wrong.
|
||
- **Update operators**: `$set` `$unset` `$inc` `$push` (`$each`) `$pull`
|
||
`$rename`, with dot-path creation (including array indices).
|
||
- **Storage**: append-only record log (CRC32-checked, `fsync` per write,
|
||
torn-tail tolerant) with in-memory indexes rebuilt on open and automatic
|
||
compaction (rewrite + atomic rename when the log grows past
|
||
`--compact-threshold`, default 16 MB). Killed mid-write (`kill -9`), the
|
||
database recovers all committed writes; the log and compaction both work
|
||
with relative or absolute `--db` paths. Records up to the announced 16 MB
|
||
`maxBsonObjectSize` replay correctly.
|
||
- **Concurrency**: a writer-preferring read/write lock splits command
|
||
execution — reads (`find`, `count`, `aggregate`, `list*`) run concurrently
|
||
across connections, writes (CRUD, DDL) are exclusive and totally ordered,
|
||
and handshake/no-op commands run lock-free. The log append + `fsync` still
|
||
happen under the write lock, so the crash guarantees are unchanged. Fine
|
||
for light workloads.
|
||
|
||
## Layout
|
||
|
||
```
|
||
src/
|
||
bson.zig BSON parse/serialize, ObjectId, canonical comparison order
|
||
wire.zig OP_MSG/OP_QUERY framing, message + reply builders
|
||
commands.zig command dispatch (hello, CRUD, aggregate, admin, indexes)
|
||
server.zig TCP accept loop, per-connection handlers, TTL sweep monitor
|
||
db.zig in-memory engine: db → collection → _id → document maps
|
||
storage.zig append-only log: records, replay, CRC validation
|
||
query.zig filter matcher, regex engine, sort, projection
|
||
index.zig secondary indexes: entries, search, query planner
|
||
update.zig update operators with dot-path navigation
|
||
main.zig CLI: --port, --bind, --db, --ttl-sweep-secs, --compact-threshold
|
||
```
|
||
|
||
## Indexes
|
||
|
||
`collection.createIndex({field: 1})` works against every driver; the index
|
||
is persisted in the log, survives restarts and compaction, and is used by
|
||
the query planner to narrow scans.
|
||
|
||
- **Key patterns**: single-field and compound (up to 32 fields), each key
|
||
`1` or `-1`. Descending order is metadata (entries are always stored
|
||
value-ascending); the default index name is MongoDB's `a_1_b_-1`.
|
||
`createIndex({_id: 1})` is an idempotent no-op — the docs map is the
|
||
`_id_` index — and `dropIndex("_id_")` errors.
|
||
- **Options**: `unique` (a conflicting write fails with E11000 naming the
|
||
index; per-document entries are deduped first, so `{a: [1,1]}` is legal)
|
||
and `sparse` (documents missing an indexed field are skipped).
|
||
- **TTL**: `createIndex({expireAt: 1}, {expireAfterSeconds: 60})` deletes a
|
||
document once its indexed date is that many seconds old. A background
|
||
sweeper runs every `--ttl-sweep-secs` seconds (default 60, `0` disables
|
||
it) and deletes through the ordinary write path, so each expiry is logged
|
||
and fsynced and holds across a restart. As in MongoDB the option is
|
||
single-field only (a compound key is `CannotCreateIndex`, code 67),
|
||
`expireAfterSeconds` must be a whole number in `[0, 2147483647]` (`0`
|
||
means "expire at the stored instant"), a non-date value at the path never
|
||
expires, an array of dates expires on its earliest member, and expiry is
|
||
coarse: a document stays visible until the next sweep. Re-creating an
|
||
index with a different expiry is `IndexOptionsConflict` (85) and an
|
||
expiry on `{_id: 1}` is `InvalidIndexSpecificationOption` (197), both as
|
||
MongoDB has them.
|
||
- **Multikey**: an array at an indexed path is indexed as a whole *and*
|
||
element-wise, mirroring the query matcher exactly, so both
|
||
`{tags: "a"}` and `{tags: ["a","b"]}` hit the index. A compound index
|
||
over two array paths rejects the document with MongoDB's "cannot index
|
||
parallel arrays".
|
||
- **Planner**: picks the index covering the longest leading run of
|
||
equality/`$in` predicates (cartesian product capped at 100 lookups),
|
||
optionally with a range on the next key. Ranges with both bounds fall
|
||
back to a scan on multikey indexes (a doc with `{a: [1,2]}` can satisfy
|
||
`{a: {$gt: 5, $lt: 25}}` across two entries), and sparse indexes are
|
||
never used for `null`-valued predicates. The `_id_` fast path resolves
|
||
`{_id: ...}` through the docs map unless the value's compare class is
|
||
serialization-ambiguous (int32 1, int64 1, double 1.0 compare equal but
|
||
hash differently — those fall back to a scan, as do string/symbol/code).
|
||
|
||
v1 limits: no index-accelerated sort, no hashed/text/geo/partial indexes,
|
||
and entry insert/removal is O(n) (a sorted array) — fine for a light
|
||
database, with a B-tree or id→entry map as the follow-up. A TTL sweep
|
||
walks every entry of every TTL index and holds the write lock for the
|
||
whole pass, so the interval is the tuning knob: expiry is never more
|
||
precise than `--ttl-sweep-secs`, and a very large TTL index wants a
|
||
longer one.
|
||
|
||
## Not (yet) implemented
|
||
|
||
- Authentication (SCRAM) — run without credentials
|
||
- Real cursors (all results are returned in one batch, cursor id 0)
|
||
- Transactions, change streams, replicasets
|
||
- Compression (OP_COMPRESSED)
|
||
- `collMod`, so an index's `expireAfterSeconds` cannot be changed in
|
||
place — drop the index and re-create it with the new expiry
|
||
- `dropCollection`/`dropDatabase` write no log record, so a dropped
|
||
collection (and its index definitions) resurrect on restart
|
||
|
||
## Working with large collections
|
||
|
||
Everything lives in RAM (db → collection → _id → document maps) and every
|
||
write command is logged with `fsync` before it is acknowledged (one sync per
|
||
command via group commit — a 500-doc `insertMany` syncs once, not 500
|
||
times), so multi-GB collections work, with cost/behavior notes measured by
|
||
the `tests/e2e/big.js` harness (12-core/32 GB Mac):
|
||
|
||
- **Build in ReleaseFast** — `zig build` defaults to it. A Debug server is
|
||
10-200x slower on every path (the matcher alone was 70 µs/doc in Debug
|
||
vs 0.4 µs in ReleaseFast), which dwarfed every other difference in the
|
||
MongoDB comparison below.
|
||
- **Compaction is O(n²) under the default 16 MB threshold.** A compaction
|
||
rewrites the whole log (one fsync per record), so bulk-loading 5 GB with
|
||
the default threshold degrades from ~310 MB/s to a crawl as the dataset
|
||
grows. Raise `--compact-threshold` for bulk loads — e.g. `2g` — and the
|
||
rate stays flat. The 5.37 GB run (40,960 × 128 KB docs, ObjectIds,
|
||
ReleaseFast) inserted in 36.5 s at ~310 MB/s between the two threshold
|
||
compactions, peaked at 5.25 GB RSS (~0.98x the data size at 128 KB
|
||
docs), and reopened the 5 GB log in 13.8 s.
|
||
- **`findOne({_id})` is O(1) only for ObjectId ids.** Integer, int64 and
|
||
double ids compare equal but hash differently, so the docs-map fast path
|
||
is skipped and every `_id` lookup becomes a full scan. Use the driver's
|
||
default ObjectIds (or a secondary index) on big collections.
|
||
- **Secondary-index entry insert is O(n)** (sorted array — see v1 limits
|
||
above), so creating an index over existing data or inserting with an
|
||
index in place is quadratic. Create indexes after the load.
|
||
|
||
## Performance vs MongoDB
|
||
|
||
`tests/e2e/compare-run.sh` runs the same driver workload (1 GB, 65,536 ×
|
||
16 KB docs, every write durable — mongo-lite fsyncs per command, mongod
|
||
runs with `j: true`) against each server and prints a side-by-side table.
|
||
With the ReleaseFast default build (MongoDB 8.3.7 on the same Mac):
|
||
|
||
| benchmark | mongo-lite | mongodb | winner |
|
||
|---|---|---|---|
|
||
| insertOne (sequential) | 0.2 ms | 4.9 ms | **mongo-lite ×24** |
|
||
| bulk insert (insertMany) | 267 MB/s | 690 MB/s | mongodb ×2.6 |
|
||
| createIndex({k: 1}) | 0.66 s | 0.08 s | mongodb ×8 |
|
||
| countDocuments({}) | 2.5 ms | 11 ms | **mongo-lite ×4** |
|
||
| findOne({_id}) | 0.6 ms | 0.7 ms | mongo-lite |
|
||
| findOne indexed | 0.7 ms | 2.6 ms | **mongo-lite ×4** |
|
||
| range-scan count | 25 ms | 13 ms | mongodb ×2 |
|
||
| sort + limit(20) | 40 ms | 2 ms | mongodb ×20 |
|
||
| aggregate $group | 12 ms | 13 ms | mongo-lite |
|
||
| updateOne({_id}) | 0.18 ms | 0.21 ms | mongo-lite |
|
||
| updateMany (65 docs) | 20 ms | 6 ms | mongodb ×3 |
|
||
| deleteOne + insert | 0.9 ms | 5 ms | **mongo-lite ×6** |
|
||
| server RSS | 2.0 GB | 1.3 GB | mongodb (×0.65) |
|
||
| kill -9 → reopen | 3.8 s | 1.3 s | mongodb |
|
||
| db on disk | 1.0 GB | 89 MB | mongodb (compressed) |
|
||
|
||
The pattern: mongo-lite wins every *latency-bound* single-op (no network of
|
||
index hops, no journal latency, in-RAM) and loses the *throughput-bound*
|
||
bulk paths and the ops MongoDB accelerates with disk indexes and
|
||
compression.
|
||
|
||
### Suggested improvements (highest impact first)
|
||
|
||
1. **Index-accelerated sort** — the worst gap (×20): `sort+limit` sorts
|
||
every document. Stream candidates in index order (the planner already
|
||
has ordered range search) and stop at `limit`. Fixes the biggest read
|
||
regression.
|
||
2. **Batch index builds** — `createIndex` inserts entries one at a time
|
||
into a sorted array (O(n²) memmoves). Sort all entries once and append
|
||
in bulk (O(n log n)); a B-tree or id→entry map removes the O(n) entry
|
||
insert on the write path too.
|
||
3. **Compress the log** — the db is 11× MongoDB's on disk because payloads
|
||
are stored raw. Snappy per record (like the wire protocol's OP_COMPRESSED)
|
||
would shrink highly-compressible workloads massively.
|
||
4. **Faster reopen** — replay is a full re-parse of every record. A
|
||
periodic checkpoint record (or a parallel replay) would cut the 3×
|
||
restart gap.
|
||
5. **Trim the write path** — bulk insert (×2.6) is now bound by per-doc
|
||
parse/serialize/map-put, not fsync. A pooled per-connection arena for
|
||
owned docs and a bulk-insert fast path would close most of the gap;
|
||
updateMany's per-doc replace-serialize is the same story.
|
||
6. **Range-scan matching (×2)** — the matcher allocates a candidates list
|
||
per field per doc; a stack buffer for the common single-field case
|
||
removes it.
|
||
|
||
Two real bugs were found and fixed while benchmarking:
|
||
|
||
- `plan_id` returned a pointer to a stack temporary (`&.{e}`) that dangled
|
||
after the frame returned — Debug tolerated it, ReleaseFast read garbage,
|
||
silently breaking every `findOne({_id: <ObjectId>})`. It now heap-copies
|
||
the lookup value and frees it.
|
||
- Multi-doc writes fsynced once per document; they now group-commit (one
|
||
fsync per command, same crash guarantees — verified by the kill -9
|
||
crash suites).
|
||
|
||
|
||
## Code style
|
||
|
||
Zig 0.16 idioms (`std.Io` threaded through everything, unmanaged
|
||
containers); user-declared functions use `snake_case` per this repo's house
|
||
style.
|