README: current numbers and remaining gaps

The performance table, the large-collection notes and the roadmap all
described the state before the optimization work.

The table is the recorded Phase 1 gate run (tests/e2e/results/phase1.txt).
Bulk insert, createIndex and reopen moved from losses to wins; the sort
row is split, since an indexed field now streams out of the index at 1.0ms
while _id still materializes.

The large-collection notes carried advice that is now wrong: compaction no
longer needs a raised --compact-threshold for bulk loads, and building an
index over existing data is no longer quadratic (inserting into one that
already exists still is). The v1 index limits no longer claim removal is
O(n) or that sort cannot use an index.

The roadmap becomes what is left, in order, with the reason each remaining
gap is structural, plus a record of what was done and the measurement that
drove it. Two more bugs join the list: compaction fsynced once per live
document, and remove never checked the compaction threshold.
This commit is contained in:
2026-08-02 20:06:31 +03:00
parent 4aaa555563
commit b3b351167f

158
README.md
View File

@@ -123,9 +123,11 @@ the query planner to narrow scans.
serialization-ambiguous (int32 1, int64 1, double 1.0 compare equal but
hash differently — those fall back to a scan, as do string/symbol/code).
v1 limits: no index-accelerated sort, no hashed/text/geo/partial indexes,
and entry insert/removal is O(n) (a sorted array) — fine for a light
database, with a B-tree or id→entry map as the follow-up. A TTL sweep
v1 limits: no hashed/text/geo/partial indexes, and entry *insert* is O(n)
(a sorted array memmoves the tail) — fine for a light database, with a
B-tree as the follow-up. Removal is no longer a scan: entry generation is
a pure function of the document, so the entries to drop are regenerated
and found by binary search. A TTL sweep
walks every entry of every TTL index and holds the write lock for the
whole pass, so the interval is the tuning knob: expiry is never more
precise than `--ttl-sweep-secs`, and a very large TTL index wants a
@@ -154,21 +156,25 @@ the `tests/e2e/big.js` harness (12-core/32 GB Mac):
10-200x slower on every path (the matcher alone was 70 µs/doc in Debug
vs 0.4 µs in ReleaseFast), which dwarfed every other difference in the
MongoDB comparison below.
- **Compaction is O(n²) under the default 16 MB threshold.** A compaction
rewrites the whole log (one fsync per record), so bulk-loading 5 GB with
the default threshold degrades from ~310 MB/s to a crawl as the dataset
grows. Raise `--compact-threshold` for bulk loads — e.g. `2g` — and the
rate stays flat. The 5.37 GB run (40,960 × 128 KB docs, ObjectIds,
ReleaseFast) inserted in 36.5 s at ~310 MB/s between the two threshold
compactions, peaked at 5.25 GB RSS (~0.98x the data size at 128 KB
docs), and reopened the 5 GB log in 13.8 s.
- **Compaction no longer needs tuning for bulk loads.** It triggers on the
share of the log that is garbage rather than on bytes appended, so a pure
insert workload — which has no garbage — is never rewritten, and a
rewrite-heavy one is reclaimed once about a fifth of the log is dead,
keeping the file near 1.25x the live data. `--compact-threshold` is now
only a floor below which small logs are left alone. (It used to fire
every 16 MB regardless, rewriting the whole log each time: quadratic
total traffic, and the reason bulk loads needed a raised threshold.)
- **`findOne({_id})` is O(1) only for ObjectId ids.** Integer, int64 and
double ids compare equal but hash differently, so the docs-map fast path
is skipped and every `_id` lookup becomes a full scan. Use the driver's
default ObjectIds (or a secondary index) on big collections.
default ObjectIds (or a secondary index) on big collections. The
order-preserving key encoding already removes the ambiguity that forces
this; lifting the restriction waits on an ordered `_id` index.
- **Secondary-index entry insert is O(n)** (sorted array — see v1 limits
above), so creating an index over existing data or inserting with an
index in place is quadratic. Create indexes after the load.
above), so inserting into a collection that already has an index is
quadratic. Building an index over existing data is not: entries are
appended unsorted and ordered once. Still cheapest to create indexes
after the load.
## Performance vs MongoDB
@@ -179,52 +185,87 @@ With the ReleaseFast default build (MongoDB 8.3.7 on the same Mac):
| benchmark | mongo-lite | mongodb | winner |
|---|---|---|---|
| insertOne (sequential) | 0.2 ms | 4.9 ms | **mongo-lite ×24** |
| bulk insert (insertMany) | 267 MB/s | 690 MB/s | mongodb ×2.6 |
| createIndex({k: 1}) | 0.66 s | 0.08 s | mongodb ×8 |
| countDocuments({}) | 2.5 ms | 11 ms | **mongo-lite ×4** |
| findOne({_id}) | 0.6 ms | 0.7 ms | mongo-lite |
| findOne indexed | 0.7 ms | 2.6 ms | **mongo-lite ×4** |
| range-scan count | 25 ms | 13 ms | mongodb ×2 |
| sort + limit(20) | 40 ms | 2 ms | mongodb ×20 |
| aggregate $group | 12 ms | 13 ms | mongo-lite |
| updateOne({_id}) | 0.18 ms | 0.21 ms | mongo-lite |
| updateMany (65 docs) | 20 ms | 6 ms | mongodb ×3 |
| deleteOne + insert | 0.9 ms | 5 ms | **mongo-lite ×6** |
| server RSS | 2.0 GB | 1.3 GB | mongodb (×0.65) |
| kill -9 → reopen | 3.8 s | 1.3 s | mongodb |
| db on disk | 1.0 GB | 89 MB | mongodb (compressed) |
| insertOne (sequential) | 0.19 ms | 5.0 ms | **mongo-lite ×26** |
| bulk insert (insertMany) | 853 MB/s | 690 MB/s | **mongo-lite ×1.2** |
| createIndex({k: 1}) | 62 ms | 78 ms | **mongo-lite** |
| countDocuments({}) | 1.5 ms | 11.5 ms | **mongo-lite ×8** |
| findOne({_id}) | 0.48 ms | 0.54 ms | mongo-lite |
| findOne indexed | 0.64 ms | 4.3 ms | **mongo-lite ×7** |
| range-scan count | 21 ms | 13 ms | mongodb ×1.7 |
| sort + limit(20), on `_id` | 4.3 ms | 2.0 ms | mongodb ×2 |
| sort + limit(20), indexed field | 1.0 ms | — | — |
| aggregate $group | 9.8 ms | 13.7 ms | **mongo-lite** |
| updateOne({_id}) | 0.16 ms | 0.19 ms | mongo-lite |
| updateMany (65 docs) | 5.5 ms | 6.1 ms | mongo-lite |
| deleteOne + insert | 0.62 ms | 4.9 ms | **mongo-lite ×8** |
| server RSS | 2.0 GB | 1.4 GB | mongodb (×0.7) |
| kill -9 → reopen | 0.8 s | 1.3 s | **mongo-lite** |
| db on disk | 1.0 GB | 93 MB | mongodb (compressed) |
The pattern: mongo-lite wins every *latency-bound* single-op (no network of
index hops, no journal latency, in-RAM) and loses the *throughput-bound*
bulk paths and the ops MongoDB accelerates with disk indexes and
compression.
The remaining losses are structural rather than incidental. Disk size is
the big one: payloads are stored raw, so the log is 11x MongoDB's
compressed files. RSS trails because every document carries its own arena.
The range-scan gap is not the matcher — it is walking 65,536 documents
that each live in a separate allocation, one pointer chase apiece. And
`sort` on `_id` still materializes candidates because nothing ordered
covers `_id` yet; the same sort on an indexed field streams straight out
of the index at 1.0 ms.
### Suggested improvements (highest impact first)
Reproduce the table with `bash tests/e2e/compare-run.sh 1g 16k`; the run
above is recorded in `tests/e2e/results/phase1.txt`.
1. **Index-accelerated sort** — the worst gap (×20): `sort+limit` sorts
every document. Stream candidates in index order (the planner already
has ordered range search) and stop at `limit`. Fixes the biggest read
regression.
2. **Batch index builds**`createIndex` inserts entries one at a time
into a sorted array (O(n²) memmoves). Sort all entries once and append
in bulk (O(n log n)); a B-tree or id→entry map removes the O(n) entry
insert on the write path too.
3. **Compress the log** — the db is 11× MongoDB's on disk because payloads
are stored raw. Snappy per record (like the wire protocol's OP_COMPRESSED)
would shrink highly-compressible workloads massively.
4. **Faster reopen** — replay is a full re-parse of every record. A
periodic checkpoint record (or a parallel replay) would cut the 3×
restart gap.
5. **Trim the write path** — bulk insert (×2.6) is now bound by per-doc
parse/serialize/map-put, not fsync. A pooled per-connection arena for
owned docs and a bulk-insert fast path would close most of the gap;
updateMany's per-doc replace-serialize is the same story.
6. **Range-scan matching (×2)** — the matcher allocates a candidates list
per field per doc; a stack buffer for the common single-field case
removes it.
### What is left (highest impact first)
Two real bugs were found and fixed while benchmarking:
1. **Compress the log** — the largest remaining gap (×11). Payloads are
stored raw. A block-framed format with an LZ4 block codec would shrink
highly compressible workloads massively; note Zig 0.16 ships zstd
decompression only, and deflate would cap writes below the current
insert rate.
2. **A B-tree over the encoded keys** — entry insert still memmoves the
tail of a sorted array, so writing into a collection that already has an
index is quadratic. A flat, `u32`-indexed node array would also be
dumpable into a checkpoint, which is what makes a fast reopen possible.
3. **An ordered `_id` index**`sort({_id: ...})` still materializes every
candidate, and integer `_id`s still scan. Both fall out of indexing the
encoded `_id`. It wants the tree first: an `_id` index updates on every
insert, and doing that against a sorted array is only cheap because
ObjectIds append at the end.
4. **Stop giving every document its own arena** — the source of both the
RSS gap and the range-scan gap. Storing canonical BSON bytes in a
per-collection slab and matching against them (parsing only the fields a
filter names) makes scans contiguous instead of a pointer chase.
5. **Decompose the global lock** — one reader/writer lock covers the whole
engine and is held across fsync, compaction and reply construction.
Per-collection locks plus cross-connection group commit are the path to
using more than one core on writes.
Done so far, with the measurement that drove each:
- **Record integrity hash CRC32 → XxHash3.** `std.hash.Crc32` is
table-driven and byte-at-a-time: 408 MB/s against XxHash3's 31 GB/s, or
38 µs versus 0.5 µs on a 16 KB document — about two thirds of the entire
bulk-insert cost. Insert 260 → 700 MB/s.
- **Compaction triggers on garbage, not on bytes written**, and syncs once
per rewrite instead of once per document. Bulk load at the default
threshold 41.6 → 703 MB/s.
- **Index builds append then sort once** instead of inserting into a sorted
array. `createIndex` over 65,536 documents 649 → 44 ms.
- **Index entries hold encoded byte keys**, so comparing them is a memcmp
rather than a walk over values in unrelated arenas.
- **Entry removal is a binary search**, not a scan of the whole index.
`updateMany` 15.4 → 5.5 ms.
- **Top-k sort selection** and an allocation-free decorate pass, plus
**index-supplied ordering** when an index already holds candidates in the
requested order. `sort+limit(20)` 40 → 4.3 ms, or 1.0 ms on an indexed
field.
- **`limit` reaches the scan**, which used to materialize the whole
collection before slicing, and `countDocuments` is answered by counting
rather than by materializing and discarding every match.
- **Matching collects candidates on the stack**, resolves operators to an
enum once per filter field rather than by string per document, and reuses
one reply arena per connection.
Several real bugs surfaced while benchmarking:
- `plan_id` returned a pointer to a stack temporary (`&.{e}`) that dangled
after the frame returned — Debug tolerated it, ReleaseFast read garbage,
@@ -233,6 +274,11 @@ Two real bugs were found and fixed while benchmarking:
- Multi-doc writes fsynced once per document; they now group-commit (one
fsync per command, same crash guarantees — verified by the kill -9
crash suites).
- Compaction fsynced once per live document, because the log it wrote into
never had deferred syncing enabled — 65,536 fsyncs to rewrite a 1 GB
collection.
- `remove` never checked the compaction threshold, so a delete-heavy
workload grew the log without bound.
## Code style