Files
MultiforaDB/tests/e2e/results/bench-20260803-223947.txt
Aleksey Shakhmatov 504179acd1 results: the M0 gates, measured
PLAN D7's six items, with the numbers and the command that reproduces each in
tests/e2e/results/m0-gates.txt. Unit tests green in both optimize modes, the
whole e2e matrix green, the spec scorecard byte-identical at 131/161/195, and
the large smoke run at the scale D7.3 asked for:

  21.47 GB collection (1,310,720 x 16 KiB)
  data file                 21.75 GB      (+1.3% over the documents)
  log after the load        2.5 MB        (checkpoints reclaim it)
  kill -9 then reopen       0.5 s         (0.5 s at 4 GB too -- flat)
  RSS after reopen          237 MB        (1.1% of the data)
  count after restart       1,310,720     last document byte-intact
  acked writes after kill   200/200

That is the milestone's claim, measured: an open costs the working set rather
than the size of the database. Before M0 the same measurement was 523 MB
resident for a 512 MB database, because recovering each document's `_id` meant
reading every document at open.

Two gates need reading rather than a tick, and m0-gates.txt says so where a
reader would otherwise take a tick for granted.

The churn gate settles at 1.65x live data (delete-heavy) to 2.47x
(update-heavy), flat, above the ~1.3x amendment A2 hoped for. Rebuild-only
reclamation cannot reach that: it needs a whole second copy of the live data
before the first can be freed. The gate existed to decide whether doc-level
free lists are needed after M0, and that is the answer.

Benchmark parity holds for every read and latency row inside the run-to-run
spread, and bulk insert regresses 24% (732 -> 555 MB/s), reproducibly across
three runs. Risk 1 as written: document bytes now reach the disk uncompressed
on top of the LZ4 log. createIndex improves 62% from the same change.

Three measurement bugs fixed while running the gates, because each would have
put a false number in the README:

  - `compare-run.sh` measured "db on disk" as `du` of the log alone against
    `du` of mongod's whole dbpath. It reported 20 MB for a 1 GB collection --
    the documents had moved to <db>.data. Honest figure, measured: 914 MB of
    allocated blocks against mongod's compressed 85 MB.
  - `big.js` counted "compaction events" as "the log shrank", which is a
    *checkpoint* now. It claimed 12 compaction rewrites during a pure insert
    load, which has no garbage to compact.
  - `big.js` labelled peak RSS "in-memory engine: docs live in RAM" and its
    summary said the collection was held "fully in RAM". Both were true of the
    engine this milestone replaced.

README: the storage section described an all-in-RAM engine; the comparison
table mixed one old run's body with three new rows; and `findOne({_id})` was
documented as a full scan for integer ids, which the ordered `_id_` index made
false (2 ms against 55 s for a scan of the same 21.5 GB collection). The table
is now best-of-three for both servers, with the measured variance stated, since
two runs of the same binary moved the sub-10 ms rows by 27-51%.
2026-08-03 23:09:43 +03:00

37 lines
1.0 KiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# multiforadb vs MongoDB benchmark
# date: 2026-08-03T19:42:00Z git: 5228ed7+dirty
# args: size=1g doc-size=16k wc=j clients=1 4 8 16 32 per-client=2000
# reproduce: bash tests/e2e/bench-run.sh 1g 16k "1 4 8 16 32"
[main]
insertOne (sequential) ×200 0.20 ms 4.9 ms
bulk insert throughput 463.4 MB/s 581.3 MB/s
docs loaded 65,536 65,536
createIndex({k: 1}) 27.9 ms 78.1 ms
countDocuments({}) 2.6 ms 12.8 ms
findOne({_id: <ObjectId>}) 0.64 ms 0.72 ms
findOne({k: 500}) (indexed) 0.59 ms 1.6 ms
find({p: {$gte,$lt}}).count() (scan) 14.1 ms 12.5 ms
find({}).sort({_id:-1}).limit(20) 1.5 ms 2.4 ms
find({}, {proj}).limit(1000) 3.6 ms 4.7 ms
aggregate $group by k 8.4 ms 13.4 ms
updateOne({_id}) ×50 0.15 ms 0.23 ms
updateMany({k: 7}, {$inc}) 1.9 ms 7.0 ms
deleteOne({_id}) + insertOne 0.63 ms 5.8 ms
node client RSS 153 MB 155 MB
[concurrency]
clients 1 8354 204 41.0x
clients 4 465 0.0x
clients 8 978 0.0x
clients 16 1877 0.0x
clients 32 3765 0.0x
[meta]
mfdb_rss_mb 1059
md_rss_mb 1203
mfdb_reopen 0.3s
md_reopen 1.3s
mfdb_disk_mb 20MB
md_disk_mb 83MB