README documents supported key patterns, unique/sparse/multikey behavior, planner rules (multikey two-bound range fallback, sparse/null bail, the _id fast-path guards), and v1 limits plus the two pre-existing issues the work surfaces (drop-collection resurrection, compact log_bytes). e2e3.js exercises createIndex/getIndexes/dropIndex/dropIndexes, the unique-constraint 11000 path, compound, sparse, and descending indexes through the official Node driver.
129 lines
6.3 KiB
Markdown
129 lines
6.3 KiB
Markdown
# mongo-light
|
|
|
|
A lightweight, embedded MongoDB-compatible document database written in
|
|
Zig 0.16. Like SQLite, it stores everything in a single file; unlike SQLite,
|
|
it speaks the MongoDB wire protocol, so real clients — `mongosh`, the Node.js
|
|
driver, PyMongo — connect over TCP and just work.
|
|
|
|
## Quick start
|
|
|
|
```sh
|
|
zig build # build the server
|
|
zig build test # run the unit test suite
|
|
|
|
zig-out/bin/mongo-light --port 27017 --db data.log
|
|
|
|
# in another terminal:
|
|
mongosh --port 27017
|
|
> db.users.insertOne({name: "alice", age: 30})
|
|
> db.users.find({age: {$gt: 25}}).toArray()
|
|
> db.users.updateOne({name: "alice"}, {$set: {vip: true}})
|
|
> db.users.deleteOne({name: "bob"})
|
|
```
|
|
|
|
## Features
|
|
|
|
- **Wire protocol**: OP_MSG (2013) plus legacy OP_QUERY/OP_REPLY (2004/2001)
|
|
for the driver handshake; hello/isMaster with `maxWireVersion: 8`, so
|
|
modern drivers (Node, Python, mongosh) connect without workarounds.
|
|
- **BSON**: full parse/serialize round-trip for all common types
|
|
(including binary, regex, timestamps, ObjectId), canonical MongoDB
|
|
comparison order for sorting and range queries.
|
|
- **CRUD**: `insert`, `find` (filter, sort, skip/limit, projection),
|
|
`update` (multi/upsert), `delete`, `findAndModify`, `count`,
|
|
`aggregate` (`$match`, `$sort`, `$skip`, `$limit`, `$project`, `$count`,
|
|
`$group` with `$sum`), plus `create`/`drop`/`listCollections`/
|
|
`listDatabases`/`dropDatabase`.
|
|
- **Query operators**: `$eq` `$ne` `$gt` `$gte` `$lt` `$lte` `$in` `$nin`
|
|
`$exists` `$regex` (hand-rolled engine: anchors, `.`, `* + ?`, character
|
|
classes, groups, alternation, `i`/`s` options) `$not` `$and` `$or` `$nor`
|
|
`$size` `$all` `$elemMatch`, with dot paths and array multikey semantics.
|
|
- **Secondary indexes**: `createIndex`/`listIndexes`/`dropIndex` via the
|
|
three driver commands, single-field and compound, with `unique` and
|
|
`sparse` options, persisted in the log and rebuilt on open (compaction
|
|
re-emits them). The query planner turns equality / `$in` / range
|
|
predicates into index lookups across `find`, `count`, `update`,
|
|
`delete`, `findAndModify`, and a leading `$match` in `aggregate`; every
|
|
candidate is re-checked against the full filter, so an index that
|
|
over-approximates is merely slow, never wrong.
|
|
- **Update operators**: `$set` `$unset` `$inc` `$push` (`$each`) `$pull`
|
|
`$rename`, with dot-path creation (including array indices).
|
|
- **Storage**: append-only record log (CRC32-checked, `fsync` per write,
|
|
torn-tail tolerant) with in-memory indexes rebuilt on open and automatic
|
|
compaction (rewrite + atomic rename when the log grows past 16 MB).
|
|
Killed mid-write (`kill -9`), the database recovers all committed writes;
|
|
the log and compaction both work with relative or absolute `--db` paths.
|
|
Records up to the announced 16 MB `maxBsonObjectSize` replay correctly.
|
|
- **Concurrency**: a writer-preferring read/write lock splits command
|
|
execution — reads (`find`, `count`, `aggregate`, `list*`) run concurrently
|
|
across connections, writes (CRUD, DDL) are exclusive and totally ordered,
|
|
and handshake/no-op commands run lock-free. The log append + `fsync` still
|
|
happen under the write lock, so the crash guarantees are unchanged. Fine
|
|
for light workloads.
|
|
|
|
## Layout
|
|
|
|
```
|
|
src/
|
|
bson.zig BSON parse/serialize, ObjectId, canonical comparison order
|
|
wire.zig OP_MSG/OP_QUERY framing, message + reply builders
|
|
commands.zig command dispatch (hello, CRUD, aggregate, admin, indexes)
|
|
server.zig TCP accept loop, per-connection handlers
|
|
db.zig in-memory engine: db → collection → _id → document maps
|
|
storage.zig append-only log: records, replay, CRC validation
|
|
query.zig filter matcher, regex engine, sort, projection
|
|
index.zig secondary indexes: entries, search, query planner
|
|
update.zig update operators with dot-path navigation
|
|
main.zig CLI: --port, --bind, --db
|
|
```
|
|
|
|
## Indexes
|
|
|
|
`collection.createIndex({field: 1})` works against every driver; the index
|
|
is persisted in the log, survives restarts and compaction, and is used by
|
|
the query planner to narrow scans.
|
|
|
|
- **Key patterns**: single-field and compound (up to 32 fields), each key
|
|
`1` or `-1`. Descending order is metadata (entries are always stored
|
|
value-ascending); the default index name is MongoDB's `a_1_b_-1`.
|
|
`createIndex({_id: 1})` is an idempotent no-op — the docs map is the
|
|
`_id_` index — and `dropIndex("_id_")` errors.
|
|
- **Options**: `unique` (a conflicting write fails with E11000 naming the
|
|
index; per-document entries are deduped first, so `{a: [1,1]}` is legal)
|
|
and `sparse` (documents missing an indexed field are skipped).
|
|
- **Multikey**: an array at an indexed path is indexed as a whole *and*
|
|
element-wise, mirroring the query matcher exactly, so both
|
|
`{tags: "a"}` and `{tags: ["a","b"]}` hit the index. A compound index
|
|
over two array paths rejects the document with MongoDB's "cannot index
|
|
parallel arrays".
|
|
- **Planner**: picks the index covering the longest leading run of
|
|
equality/`$in` predicates (cartesian product capped at 100 lookups),
|
|
optionally with a range on the next key. Ranges with both bounds fall
|
|
back to a scan on multikey indexes (a doc with `{a: [1,2]}` can satisfy
|
|
`{a: {$gt: 5, $lt: 25}}` across two entries), and sparse indexes are
|
|
never used for `null`-valued predicates. The `_id_` fast path resolves
|
|
`{_id: ...}` through the docs map unless the value's compare class is
|
|
serialization-ambiguous (int32 1, int64 1, double 1.0 compare equal but
|
|
hash differently — those fall back to a scan, as do string/symbol/code).
|
|
|
|
v1 limits: no index-accelerated sort, no hashed/text/geo/TTL/partial
|
|
indexes, and entry insert/removal is O(n) (a sorted array) — fine for a
|
|
light database, with a B-tree or id→entry map as the follow-up.
|
|
|
|
## Not (yet) implemented
|
|
|
|
- Authentication (SCRAM) — run without credentials
|
|
- Real cursors (all results are returned in one batch, cursor id 0)
|
|
- Transactions, change streams, replicasets
|
|
- Compression (OP_COMPRESSED)
|
|
- `dropCollection`/`dropDatabase` write no log record, so a dropped
|
|
collection (and its index definitions) resurrect on restart; and
|
|
compaction never resets `log_bytes`, so every write after the first
|
|
compaction re-triggers the threshold check
|
|
|
|
## Code style
|
|
|
|
Zig 0.16 idioms (`std.Io` threaded through everything, unmanaged
|
|
containers); user-declared functions use `snake_case` per this repo's house
|
|
style.
|