Advertising topologyVersion in the hello reply is what tells a driver the server speaks the streaming (awaitable) hello protocol — in the Node driver it is the only condition checked. From the second heartbeat on, the driver then monitored with an exhaust hello (exhaustAllowed + maxAwaitTimeMS) and waited for a stream of replies carrying moreToCome. We answered once with the flag clear and went back to reading, so every heartbeat failed with "Server ended moreToCome unexpectedly", destroying the connection and clearing the pool. MongoDB Compass showed this as a connect/disconnect loop once per heartbeat. We do not implement streaming hello, so we must not claim to. Omitting the field keeps monitoring on the polling path, and agrees with the maxWireVersion 8 we report: streaming hello arrived in wire version 9. The existing e2e files all passed against the broken server — they issue their commands and exit before the second heartbeat — so e2e5 watches SDAM heartbeats on an idle connection instead. Also renames mongo-light to mongo-lite throughout (binary, log messages, docs, gitVersion). Unrelated to the fix above, but squashed in at request rather than left as a commit whose message described only the fix.
151 lines
7.7 KiB
Markdown
151 lines
7.7 KiB
Markdown
# mongo-lite
|
|
|
|
A lightweight, embedded MongoDB-compatible document database written in
|
|
Zig 0.16. Like SQLite, it stores everything in a single file; unlike SQLite,
|
|
it speaks the MongoDB wire protocol, so real clients — `mongosh`, the Node.js
|
|
driver, PyMongo — connect over TCP and just work.
|
|
|
|
## Quick start
|
|
|
|
```sh
|
|
zig build # build the server
|
|
zig build test # run the unit test suite
|
|
|
|
zig-out/bin/mongo-lite --port 27017 --db data.log
|
|
|
|
# in another terminal:
|
|
mongosh --port 27017
|
|
> db.users.insertOne({name: "alice", age: 30})
|
|
> db.users.find({age: {$gt: 25}}).toArray()
|
|
> db.users.updateOne({name: "alice"}, {$set: {vip: true}})
|
|
> db.users.deleteOne({name: "bob"})
|
|
> db.sessions.createIndex({expireAt: 1}, {expireAfterSeconds: 3600})
|
|
```
|
|
|
|
## Features
|
|
|
|
- **Wire protocol**: OP_MSG (2013) plus legacy OP_QUERY/OP_REPLY (2004/2001)
|
|
for the driver handshake; hello/isMaster with `maxWireVersion: 8`, so
|
|
modern drivers (Node, Python, mongosh) connect without workarounds.
|
|
- **BSON**: full parse/serialize round-trip for all common types
|
|
(including binary, regex, timestamps, ObjectId), canonical MongoDB
|
|
comparison order for sorting and range queries.
|
|
- **CRUD**: `insert`, `find` (filter, sort, skip/limit, projection),
|
|
`update` (multi/upsert), `delete`, `findAndModify`, `count`,
|
|
`aggregate` (`$match`, `$sort`, `$skip`, `$limit`, `$project`, `$count`,
|
|
`$group` with `$sum`), plus `create`/`drop`/`listCollections`/
|
|
`listDatabases`/`dropDatabase`.
|
|
- **Query operators**: `$eq` `$ne` `$gt` `$gte` `$lt` `$lte` `$in` `$nin`
|
|
`$exists` `$regex` (hand-rolled engine: anchors, `.`, `* + ?`, character
|
|
classes, groups, alternation, `i`/`s` options) `$not` `$and` `$or` `$nor`
|
|
`$size` `$all` `$elemMatch`, with dot paths and array multikey semantics.
|
|
- **Secondary indexes**: `createIndex`/`listIndexes`/`dropIndex` via the
|
|
three driver commands, single-field and compound, with `unique`,
|
|
`sparse` and `expireAfterSeconds` (TTL) options, persisted in the log
|
|
and rebuilt on open (compaction re-emits them). A background sweeper
|
|
expires TTL-indexed documents through the ordinary logged write path.
|
|
The query planner turns equality / `$in` / range
|
|
predicates into index lookups across `find`, `count`, `update`,
|
|
`delete`, `findAndModify`, and a leading `$match` in `aggregate`; every
|
|
candidate is re-checked against the full filter, so an index that
|
|
over-approximates is merely slow, never wrong.
|
|
- **Update operators**: `$set` `$unset` `$inc` `$push` (`$each`) `$pull`
|
|
`$rename`, with dot-path creation (including array indices).
|
|
- **Storage**: append-only record log (CRC32-checked, `fsync` per write,
|
|
torn-tail tolerant) with in-memory indexes rebuilt on open and automatic
|
|
compaction (rewrite + atomic rename when the log grows past 16 MB).
|
|
Killed mid-write (`kill -9`), the database recovers all committed writes;
|
|
the log and compaction both work with relative or absolute `--db` paths.
|
|
Records up to the announced 16 MB `maxBsonObjectSize` replay correctly.
|
|
- **Concurrency**: a writer-preferring read/write lock splits command
|
|
execution — reads (`find`, `count`, `aggregate`, `list*`) run concurrently
|
|
across connections, writes (CRUD, DDL) are exclusive and totally ordered,
|
|
and handshake/no-op commands run lock-free. The log append + `fsync` still
|
|
happen under the write lock, so the crash guarantees are unchanged. Fine
|
|
for light workloads.
|
|
|
|
## Layout
|
|
|
|
```
|
|
src/
|
|
bson.zig BSON parse/serialize, ObjectId, canonical comparison order
|
|
wire.zig OP_MSG/OP_QUERY framing, message + reply builders
|
|
commands.zig command dispatch (hello, CRUD, aggregate, admin, indexes)
|
|
server.zig TCP accept loop, per-connection handlers, TTL sweep monitor
|
|
db.zig in-memory engine: db → collection → _id → document maps
|
|
storage.zig append-only log: records, replay, CRC validation
|
|
query.zig filter matcher, regex engine, sort, projection
|
|
index.zig secondary indexes: entries, search, query planner
|
|
update.zig update operators with dot-path navigation
|
|
main.zig CLI: --port, --bind, --db, --ttl-sweep-secs
|
|
```
|
|
|
|
## Indexes
|
|
|
|
`collection.createIndex({field: 1})` works against every driver; the index
|
|
is persisted in the log, survives restarts and compaction, and is used by
|
|
the query planner to narrow scans.
|
|
|
|
- **Key patterns**: single-field and compound (up to 32 fields), each key
|
|
`1` or `-1`. Descending order is metadata (entries are always stored
|
|
value-ascending); the default index name is MongoDB's `a_1_b_-1`.
|
|
`createIndex({_id: 1})` is an idempotent no-op — the docs map is the
|
|
`_id_` index — and `dropIndex("_id_")` errors.
|
|
- **Options**: `unique` (a conflicting write fails with E11000 naming the
|
|
index; per-document entries are deduped first, so `{a: [1,1]}` is legal)
|
|
and `sparse` (documents missing an indexed field are skipped).
|
|
- **TTL**: `createIndex({expireAt: 1}, {expireAfterSeconds: 60})` deletes a
|
|
document once its indexed date is that many seconds old. A background
|
|
sweeper runs every `--ttl-sweep-secs` seconds (default 60, `0` disables
|
|
it) and deletes through the ordinary write path, so each expiry is logged
|
|
and fsynced and holds across a restart. As in MongoDB the option is
|
|
single-field only (a compound key is `CannotCreateIndex`, code 67),
|
|
`expireAfterSeconds` must be a whole number in `[0, 2147483647]` (`0`
|
|
means "expire at the stored instant"), a non-date value at the path never
|
|
expires, an array of dates expires on its earliest member, and expiry is
|
|
coarse: a document stays visible until the next sweep. Re-creating an
|
|
index with a different expiry is `IndexOptionsConflict` (85) and an
|
|
expiry on `{_id: 1}` is `InvalidIndexSpecificationOption` (197), both as
|
|
MongoDB has them.
|
|
- **Multikey**: an array at an indexed path is indexed as a whole *and*
|
|
element-wise, mirroring the query matcher exactly, so both
|
|
`{tags: "a"}` and `{tags: ["a","b"]}` hit the index. A compound index
|
|
over two array paths rejects the document with MongoDB's "cannot index
|
|
parallel arrays".
|
|
- **Planner**: picks the index covering the longest leading run of
|
|
equality/`$in` predicates (cartesian product capped at 100 lookups),
|
|
optionally with a range on the next key. Ranges with both bounds fall
|
|
back to a scan on multikey indexes (a doc with `{a: [1,2]}` can satisfy
|
|
`{a: {$gt: 5, $lt: 25}}` across two entries), and sparse indexes are
|
|
never used for `null`-valued predicates. The `_id_` fast path resolves
|
|
`{_id: ...}` through the docs map unless the value's compare class is
|
|
serialization-ambiguous (int32 1, int64 1, double 1.0 compare equal but
|
|
hash differently — those fall back to a scan, as do string/symbol/code).
|
|
|
|
v1 limits: no index-accelerated sort, no hashed/text/geo/partial indexes,
|
|
and entry insert/removal is O(n) (a sorted array) — fine for a light
|
|
database, with a B-tree or id→entry map as the follow-up. A TTL sweep
|
|
walks every entry of every TTL index and holds the write lock for the
|
|
whole pass, so the interval is the tuning knob: expiry is never more
|
|
precise than `--ttl-sweep-secs`, and a very large TTL index wants a
|
|
longer one.
|
|
|
|
## Not (yet) implemented
|
|
|
|
- Authentication (SCRAM) — run without credentials
|
|
- Real cursors (all results are returned in one batch, cursor id 0)
|
|
- Transactions, change streams, replicasets
|
|
- Compression (OP_COMPRESSED)
|
|
- `collMod`, so an index's `expireAfterSeconds` cannot be changed in
|
|
place — drop the index and re-create it with the new expiry
|
|
- `dropCollection`/`dropDatabase` write no log record, so a dropped
|
|
collection (and its index definitions) resurrect on restart; and
|
|
compaction never resets `log_bytes`, so every write after the first
|
|
compaction re-triggers the threshold check
|
|
|
|
## Code style
|
|
|
|
Zig 0.16 idioms (`std.Io` threaded through everything, unmanaged
|
|
containers); user-declared functions use `snake_case` per this repo's house
|
|
style.
|