style: adopt TigerStyle across src/; add docs/TIGER_STYLE.md

Wrap signatures and long expressions to the 100-column limit and make every
file zig fmt clean. Semantics-preserving throughout: ignoring whitespace and
the trailing commas that wrapping introduces, every file here is byte-identical
to its predecessor, and the one apparent exception is a warning string split
with `++`, which concatenates at comptime to the same bytes.

src/index.zig and src/commands.zig are reformatted in the commits that follow,
because their reformat is interleaved with in-flight changes to them and
separating the two would need the reformat re-derived rather than moved.
This commit is contained in:
2026-08-03 17:08:21 +03:00
parent d4c9b04f21
commit 86ae8fa8af
14 changed files with 1130 additions and 126 deletions

View File

@@ -5,6 +5,13 @@ Zig 0.16. Like SQLite, it stores everything in a single file; unlike SQLite,
it speaks the MongoDB wire protocol, so real clients — `mongosh`, the Node.js it speaks the MongoDB wire protocol, so real clients — `mongosh`, the Node.js
driver, PyMongo — connect over TCP and just work. driver, PyMongo — connect over TCP and just work.
## Forward plan
The direction from this MVP — a full-fledged embedded, tens-of-GB,
maximally MongoDB-compatible database — its decision record, milestones
and gates live in [PLAN.md](PLAN.md). Milestone 0 (mmap + WAL storage
foundation) is next.
## Quick start ## Quick start
```sh ```sh
@@ -317,6 +324,9 @@ Several real bugs surfaced while benchmarking:
## Code style ## Code style
Zig 0.16 idioms (`std.Io` threaded through everything, unmanaged The project follows TigerBeetle's TigerStyle — see
containers); user-declared functions use `snake_case` per this repo's house [`docs/TIGER_STYLE.md`](docs/TIGER_STYLE.md) (binding reference) and the
style. "Code style" section of `AGENTS.md` for the project-specific rules and
deliberate deviations. Highlights: `zig fmt` clean, 100-column hard limit,
4-space indent, snake_case, functions under 70 lines, always-on
assertions via `src/assert.zig`.

516
docs/TIGER_STYLE.md Normal file
View File

@@ -0,0 +1,516 @@
<!-- Adopted by MultiforaDB. Source: https://github.com/tigerbeetle/tigerbeetle/blob/main/docs/TIGER_STYLE.md
(TigerBeetle, Apache-2.0 / MIT). Body below is kept verbatim; it is the
project's binding style reference. Where this repo deviates, the
deviations are documented in AGENTS.md ("Code style"). -->
# TigerStyle
## The Essence Of Style
> “There are three things extremely hard: steel, a diamond, and to know one's self.” — Benjamin
> Franklin
TigerBeetle's coding style is evolving. A collective give-and-take at the intersection of
engineering and art. Numbers and human intuition. Reason and experience. First principles and
knowledge. Precision and poetry. Just like music. A tight beat. A rare groove. Words that rhyme and
rhymes that break. Biodigital jazz. This is what we've learned along the way. The best is yet to
come.
## Why Have Style?
Another word for style is design.
> “The design is not just what it looks like and feels like. The design is how it works.” — Steve
> Jobs
Our design goals are safety, performance, and developer experience. In that order. All three are
important. Good style advances these goals. Does the code make for more or less safety, performance
or developer experience? That is why we need style.
Put this way, style is more than readability, and readability is table stakes, a means to an end
rather than an end in itself.
> “...in programming, style is not something to pursue directly. Style is necessary only where
> understanding is missing.” ─ [Let Over
> Lambda](https://letoverlambda.com/index.cl/guest/chap1.html)
This document explores how we apply these design goals to coding style. First, a word on simplicity,
elegance and technical debt.
## On Simplicity And Elegance
Simplicity is not a free pass. It's not in conflict with our design goals. It need not be a
concession or a compromise.
Rather, simplicity is how we bring our design goals together, how we identify the “super idea” that
solves the axes simultaneously, to achieve something elegant.
> “Simplicity and elegance are unpopular because they require hard work and discipline to achieve” —
> Edsger Dijkstra
Contrary to popular belief, simplicity is also not the first attempt but the hardest revision. It's
easy to say “let's do something simple”, but to do that in practice takes thought, multiple passes,
many sketches, and still we may have to [“throw one
away”](https://en.wikipedia.org/wiki/The_Mythical_Man-Month).
The hardest part, then, is how much thought goes into everything.
We spend this mental energy upfront, proactively rather than reactively, because we know that when
the thinking is done, what is spent on the design will be dwarfed by the implementation and testing,
and then again by the costs of operation and maintenance.
An hour or day of design is worth weeks or months in production:
> “the simple and elegant systems tend to be easier and faster to design and get right, more
> efficient in execution, and much more reliable” — Edsger Dijkstra
## Technical Debt
What could go wrong? What's wrong? Which question would we rather ask? The former, because code,
like steel, is less expensive to change while it's hot. A problem solved in production is many times
more expensive than a problem solved in implementation, or a problem solved in design.
Since it's hard enough to discover showstoppers, when we do find them, we solve them. We don't allow
potential memcpy latency spikes, or exponential complexity algorithms to slip through.
> “You shall not pass!” — Gandalf
In other words, TigerBeetle has a “zero technical debt” policy. We do it right the first time. This
is important because the second time may not transpire, and because doing good work, that we can be
proud of, builds momentum.
We know that what we ship is solid. We may lack crucial features, but what we have meets our design
goals. This is the only way to make steady incremental progress, knowing that the progress we have
made is indeed progress.
## Safety
> “The rules act like the seat-belt in your car: initially they are perhaps a little uncomfortable,
> but after a while their use becomes second-nature and not using them becomes unimaginable.” —
> Gerard J. Holzmann
[NASA's Power of Ten — Rules for Developing Safety Critical
Code](https://spinroot.com/gerard/pdf/P10.pdf) will change the way you code forever. To expand:
- Use **only very simple, explicit control flow** for clarity. **Do not use recursion** to ensure
that all executions that should be bounded are bounded. Use **only a minimum of excellent
abstractions** but only if they make the best sense of the domain. Abstractions are [never zero
cost](https://isaacfreund.com/blog/2022-05/). Every abstraction introduces the risk of a leaky
abstraction.
- **Put a limit on everything** because, in reality, this is what we expect—everything has a limit.
For example, all loops and all queues must have a fixed upper bound to prevent infinite loops or
tail latency spikes. This follows the [“fail-fast”](https://en.wikipedia.org/wiki/Fail-fast)
principle so that violations are detected sooner rather than later. Where a loop cannot terminate
(e.g. an event loop), this must be asserted.
- Use explicitly-sized types like `u32` for everything, avoid architecture-specific `usize`.
- **Assertions detect programmer errors. Unlike operating errors, which are expected and which must
be handled, assertion failures are unexpected. The only correct way to handle corrupt code is to
crash. Assertions downgrade catastrophic correctness bugs into liveness bugs. Assertions are a
force multiplier for discovering bugs by fuzzing.**
- **Assert all function arguments and return values, pre/postconditions and invariants.** A
function must not operate blindly on data it has not checked. The purpose of a function is to
increase the probability that a program is correct. Assertions within a function are part of how
functions serve this purpose. The assertion density of the code must average a minimum of two
assertions per function.
- **[Pair assertions](https://tigerbeetle.com/blog/2023-12-27-it-takes-two-to-contract).** For
every property you want to enforce, try to find at least two different code paths where an
assertion can be added. For example, assert validity of data right before writing it to disk,
and also immediately after reading from disk.
- On occasion, you may use a blatantly true assertion instead of a comment as stronger
documentation where the assertion condition is critical and surprising.
- Split compound assertions: prefer `assert(a); assert(b);` over `assert(a and b);`.
The former is simpler to read, and provides more precise information if the condition fails.
- Use single-line `if` to assert an implication: `if (a) assert(b)`.
- **Assert the relationships of compile-time constants** as a sanity check, and also to document
and enforce [subtle
invariants](https://github.com/coilhq/tigerbeetle/blob/db789acfb93584e5cb9f331f9d6092ef90b53ea6/src/vsr/journal.zig#L45-L47)
or [type
sizes](https://github.com/coilhq/tigerbeetle/blob/578ac603326e1d3d33532701cb9285d5d2532fe7/src/ewah.zig#L41-L53).
Compile-time assertions are extremely powerful because they are able to check a program's design
integrity _before_ the program even executes.
- **The golden rule of assertions is to assert the _positive space_ that you do expect AND to
assert the _negative space_ that you do not expect** because where data moves across the
valid/invalid boundary between these spaces is where interesting bugs are often found. This is
also why **tests must test exhaustively**, not only with valid data but also with invalid data,
and as valid data becomes invalid.
- Assertions are a safety net, not a substitute for human understanding. With simulation testing,
there is the temptation to trust the fuzzer. But a fuzzer can prove only the presence of bugs,
not their absence. Therefore:
- Build a precise mental model of the code first,
- encode your understanding in the form of assertions,
- write the code and comments to explain and justify the mental model to your reviewer,
- and use VOPR as the final line of defense, to find bugs in your and reviewer's understanding
of code.
- All memory must be statically allocated at startup. **No memory may be dynamically allocated (or
freed and reallocated) after initialization.** This avoids unpredictable behavior that can
significantly affect performance, and avoids use-after-free. As a second-order effect, it is our
experience that this also makes for more efficient, simpler designs that are more performant and
easier to maintain and reason about, compared to designs that do not consider all possible memory
usage patterns upfront as part of the design.
- Declare variables at the **smallest possible scope**, and **minimize the number of variables in
scope**, to reduce the probability that variables are misused.
- There's a sharp discontinuity between a function fitting on a screen, and having to scroll to
see how long it is. For this physical reason we enforce a **hard limit of 70 lines per function**.
Art is born of constraints. There are many ways to cut a wall of code into chunks of 70 lines,
but only a few splits will feel right. Some rules of thumb:
* Good function shape is often the inverse of an hourglass: a few parameters, a simple return
type, and a lot of meaty logic between the braces.
* Centralize control flow. When splitting a large function, try to keep all switch/if
statements in the "parent" function, and move non-branchy logic fragments to helper
functions. Divide responsibility. All control flow should be handled by _one_ function, the rest shouldn't
care about control flow at all. In other words,
["push `if`s up and `for`s down"](https://matklad.github.io/2023/11/15/push-ifs-up-and-fors-down.html).
* Similarly, centralize state manipulation. Let the parent function keep all relevant state in
local variables, and use helpers to compute what needs to change, rather than applying the
change directly. Keep leaf functions pure.
- Appreciate, from day one, **all compiler warnings at the compiler's strictest setting**.
- Whenever your program has to interact with external entities, **don't do things directly in
reaction to external events**. Instead, your program should run at its own pace. Not only does
this make your program safer by keeping the control flow of your program under your control, it
also improves performance for the same reason (you get to batch, instead of context switching on
every event). Additionally, this makes it easier to maintain bounds on work done per time period.
Beyond these rules:
- Compound conditions that evaluate multiple booleans make it difficult for the reader to verify
that all cases are handled. Split compound conditions into simple conditions using nested
`if/else` branches. Split complex `else if` chains into `else { if { } }` trees. This makes the
branches and cases clear. Again, consider whether a single `if` does not also need a matching
`else` branch, to ensure that the positive and negative spaces are handled or asserted.
- Negations are not easy! State invariants positively. When working with lengths and indexes, this
form is easy to get right (and understand):
```zig
if (index < length) {
// The invariant holds.
} else {
// The invariant doesn't hold.
}
```
This form is harder, and also goes against the grain of how `index` would typically be compared to
`length`, for example, in a loop condition:
```zig
if (index >= length) {
// It's not true that the invariant holds.
}
```
- All errors must be handled. An [analysis of production failures in distributed data-intensive
systems](https://www.usenix.org/system/files/conference/osdi14/osdi14-paper-yuan.pdf) found that
the majority of catastrophic failures could have been prevented by simple testing of error
handling code.
> “Specifically, we found that almost all (92%) of the catastrophic system failures are the result
> of incorrect handling of non-fatal errors explicitly signaled in software.”
- **Always motivate, always say why**. Never forget to say why. Because if you explain the rationale
for a decision, it not only increases the hearer's understanding, and makes them more likely to
adhere or comply, but it also shares criteria with them with which to evaluate the decision and
its importance.
- **Explicitly pass options to library functions at the call site, instead of relying on the
defaults**. For example, write `@prefetch(a, .{ .cache = .data, .rw = .read, .locality = 3 });`
over `@prefetch(a, .{});`. This improves readability but most of all avoids latent, potentially
catastrophic bugs in case the library ever changes its defaults.
## Performance
> “The lack of back-of-the-envelope performance sketches is the root of all evil.” — Rivacindela
> Hudsoni
- Think about performance from the outset, from the beginning. **The best time to solve performance,
to get the huge 1000x wins, is in the design phase, which is precisely when we can't measure or
profile.** It's also typically harder to fix a system after implementation and profiling, and the
gains are less. So you have to have mechanical sympathy. Like a carpenter, work with the grain.
- **Perform back-of-the-envelope sketches with respect to the four resources (network, disk, memory,
CPU) and their two main characteristics (bandwidth, latency).** Sketches are cheap. Use sketches
to be “roughly right” and land within 90% of the global maximum.
- Optimize for the slowest resources first (network, disk, memory, CPU) in that order, after
compensating for the frequency of usage, because faster resources may be used many times more. For
example, a memory cache miss may be as expensive as a disk fsync, if it happens many times more.
- Distinguish between the control plane and data plane. A clear delineation between control plane
and data plane through the use of batching enables a high level of assertion safety without losing
performance. See our [July 2021 talk on Zig SHOWTIME](https://youtu.be/BH2jvJ74npM?t=1958) for
examples.
- Amortize network, disk, memory and CPU costs by batching accesses.
- Let the CPU be a sprinter doing the 100m. Be predictable. Don't force the CPU to zig zag and
change lanes. Give the CPU large enough chunks of work. This comes back to batching.
- Be explicit. Minimize dependence on the compiler to do the right thing for you.
In particular, extract hot loops into stand-alone functions with primitive arguments without
`self` (see [an example](https://github.com/tigerbeetle/tigerbeetle/blob/0.16.19/src/lsm/compaction.zig#L1932-L1937)).
That way, the compiler doesn't need to prove that it can cache struct's fields in registers, and a
human reader can spot redundant computations easier.
## Developer Experience
> “There are only two hard things in Computer Science: cache invalidation, naming things, and
> off-by-one errors.” — Phil Karlton
### Naming Things
- **Get the nouns and verbs just right.** Great names are the essence of great code, they capture
what a thing is or does, and provide a crisp, intuitive mental model. They show that you
understand the domain. Take time to find the perfect name, to find nouns and verbs that work
together, so that the whole is greater than the sum of its parts.
- Use `snake_case` for function, variable, and file names. The underscore is the closest thing we
have as programmers to a space, and helps to separate words and encourage descriptive names. We
don't use Zig's `CamelCase.zig` style for "struct" files to keep the convention simple and
consistent.
- Do not abbreviate variable names, unless the variable is a primitive integer type used as an
argument to a sort function or matrix calculation. Use long form arguments in scripts: `--force`,
not `-f`. Single letter flags are for interactive usage.
- Use proper capitalization for acronyms (`VSRState`, not `VsrState`).
- For the rest, follow the Zig style guide.
- Add units or qualifiers to variable names, and put the units or qualifiers last, sorted by
descending significance, so that the variable starts with the most significant word, and ends with
the least significant word. For example, `latency_ms_max` rather than `max_latency_ms`. This will
then line up nicely when `latency_ms_min` is added, as well as group all variables that relate to
latency.
- Infuse names with meaning. For example, `allocator: Allocator` is a good, if boring name,
but `gpa: Allocator` and `arena: Allocator` are excellent. They inform the reader whether
`deinit` should be called explicitly.
- When choosing related names, try hard to find names with the same number of characters so that
related variables all line up in the source. For example, as arguments to a memcpy function,
`source` and `target` are better than `src` and `dest` because they have the second-order effect
that any related variables such as `source_offset` and `target_offset` will all line up in
calculations and slices. This makes the code symmetrical, with clean blocks that are easier for
the eye to parse and for the reader to check.
- When a single function calls out to a helper function or callback, prefix the name of the helper
function with the name of the calling function to show the call history. For example,
`read_sector()` and `read_sector_callback()`.
- Callbacks go last in the list of parameters. This mirrors control flow: callbacks are also
_invoked_ last.
- _Order_ matters for readability (even if it doesn't affect semantics). On the first read, a file
is read top-down, so put important things near the top. The `main` function goes first.
The same goes for `structs`, the order is fields then types then methods:
```zig
time: Time,
process_id: ProcessID,
const ProcessID = struct { cluster: u128, replica: u8 };
const Tracer = @This(); // This alias concludes the types section.
pub fn init(gpa: std.mem.Allocator, time: Time) !Tracer {
...
}
```
If a nested type is complex, make it a top-level struct.
At the same time, not everything has a single right order. When in doubt, consider sorting
alphabetically, taking advantage of big-endian naming.
- Don't overload names with multiple meanings that are context-dependent. For example, TigerBeetle
has a feature called _pending transfers_ where a pending transfer can be subsequently _posted_ or
_voided_. At first, we called them _two-phase commit transfers_, but this overloaded the
_two-phase commit_ terminology that was used in our consensus protocol, causing confusion.
- Think of how names will be used outside the code, in documentation or communication. For example,
a noun is often a better descriptor than an adjective or present participle, because a noun can be
directly used in correspondence without having to be rephrased. Compare `replica.pipeline` vs
`replica.preparing`. The former can be used directly as a section header in a document or
conversation, whereas the latter must be clarified. Noun names compose more clearly for derived
identifiers, e.g. `config.pipeline_max`.
- Zig has named arguments through the `options: struct` pattern. Use it when arguments can be
mixed up. A function taking two `u64` must use an options struct. If an argument can be `null`,
it should be named so that the meaning of `null` literal at the call site is clear.
Because dependencies like an allocator or a tracer are singletons with unique types, they should
be threaded through constructors positionally, from the most general to the most specific.
- **Write descriptive commit messages** that inform and delight the reader, because your commit
messages are being read. Note that a pull request description is not stored in the git repository
and is invisible in `git blame`, and therefore is not a replacement for a commit message.
- Don't forget to say why. Code alone is not documentation. Use comments to explain why you wrote
the code the way you did. Show your workings.
- Don't forget to say how. For example, when writing a test, think of writing a description at the
top to explain the goal and methodology of the test, to help your reader get up to speed, or to
skip over sections, without forcing them to dive in.
- Comments are sentences, with a space after the slash, with a capital letter and a full stop, or a
colon if they relate to something that follows. Comments are well-written prose describing the
code, not just scribblings in the margin. Comments after the end of a line _can_ be phrases, with
no punctuation.
### Cache Invalidation
- Don't duplicate variables or take aliases to them. This will reduce the probability that state
gets out of sync.
- If you don't mean a function argument to be copied when passed by value, and if the argument type
is more than 16 bytes, then pass the argument as `*const`. This will catch bugs where the caller
makes an accidental copy on the stack before calling the function.
- Construct larger structs _in-place_ by passing an _out pointer_ during initialization.
In-place initializations can assume **pointer stability** and **immovable types** while
eliminating intermediate copy-move allocations, which can lead to undesirable stack growth.
Keep in mind that in-place initializations are viral — if any field is initialized
in-place, the entire container struct should be initialized in-place as well.
**Prefer:**
```zig
fn init(target: *LargeStruct) !void {
target.* = .{
// in-place initialization.
};
}
fn main() !void {
var target: LargeStruct = undefined;
try target.init();
}
```
**Over:**
```zig
fn init() !LargeStruct {
return LargeStruct {
// moving the initialized object.
}
}
fn main() !void {
var target = try LargeStruct.init();
}
```
- **Shrink the scope** to minimize the number of variables at play and reduce the probability that
the wrong variable is used.
- Calculate or check variables close to where/when they are used. **Don't introduce variables before
they are needed.** Don't leave them around where they are not. This will reduce the probability of
a POCPOU (place-of-check to place-of-use), a distant cousin to the infamous
[TOCTOU](https://en.wikipedia.org/wiki/Time-of-check_to_time-of-use). Most bugs come down to a
semantic gap, caused by a gap in time or space, because it's harder to check code that's not
contained along those dimensions.
- Use simpler function signatures and return types to reduce dimensionality at the call site, the
number of branches that need to be handled at the call site, because this dimensionality can also
be viral, propagating through the call chain. For example, as a return type, `void` trumps `bool`,
`bool` trumps `u64`, `u64` trumps `?u64`, and `?u64` trumps `!u64`.
- Ensure that functions run to completion without suspending, so that precondition assertions are
true throughout the lifetime of the function. These assertions are useful documentation without a
suspend, but may be misleading otherwise.
- Be on your guard for **[buffer bleeds](https://en.wikipedia.org/wiki/Heartbleed)**. This is a
buffer underflow, the opposite of a buffer overflow, where a buffer is not fully utilized, with
padding not zeroed correctly. This may not only leak sensitive information, but may cause
deterministic guarantees as required by TigerBeetle to be violated.
- Use newlines to **group resource allocation and deallocation**, i.e. before the resource
allocation and after the corresponding `defer` statement, to make leaks easier to spot.
### Off-By-One Errors
- **The usual suspects for off-by-one errors are casual interactions between an `index`, a `count`
or a `size`.** These are all primitive integer types, but should be seen as distinct types, with
clear rules to cast between them. To go from an `index` to a `count` you need to add one, since
indexes are _0-based_ but counts are _1-based_. To go from a `count` to a `size` you need to
multiply by the unit. Again, this is why including units and qualifiers in variable names is
important.
- Show your intent with respect to division. For example, use `@divExact()`, `@divFloor()` or
`div_ceil()` to show the reader you've thought through all the interesting scenarios where
rounding may be involved.
### Style By The Numbers
- Run `zig fmt`.
- Use 4 spaces of indentation, rather than 2 spaces, as that is more obvious to the eye at a
distance.
- Hard limit all line lengths, without exception, to at most 100 columns for a good typographic
"measure". Use it up. Never go beyond. Nothing should be hidden by a horizontal scrollbar. Let
your editor help you by setting a column ruler. To wrap a function signature, call or data
structure, add a trailing comma, close your eyes and let `zig fmt` do the rest.
Similar to function length, the motivation behind the number 100 is physical: just enough
to fit two copies of the code side-by-side on a screen.
- Add braces to the `if` statement unless it fits on a single line for consistency and defense in
depth against "goto fail;" bugs.
### Dependencies
TigerBeetle has **a “zero dependencies” policy**, apart from the Zig toolchain. Dependencies, in
general, inevitably lead to supply chain attacks, safety and performance risk, and slow install
times. For foundational infrastructure in particular, the cost of any dependency is further
amplified throughout the rest of the stack.
### Tooling
Similarly, tools have costs. A small standardized toolbox is simpler to operate than an array of
specialized instruments each with a dedicated manual. Our primary tool is Zig. It may not be the
best for everything, but it's good enough for most things. We invest into our Zig tooling to ensure
that we can tackle new problems quickly, with a minimum of accidental complexity in our local
development environment.
> “The right tool for the job is often the tool you are already using—adding new tools has a higher
> cost than many people appreciate” — John Carmack
For example, the next time you write a script, instead of `scripts/*.sh`, write `scripts/*.zig`.
This not only makes your script cross-platform and portable, but introduces type safety and
increases the probability that running your script will succeed for everyone on the team, instead of
hitting a Bash/Shell/OS-specific issue.
Standardizing on Zig for tooling is important to ensure that we reduce dimensionality, as the team,
and therefore the range of personal tastes, grows. This may be slower for you in the short term, but
makes for more velocity for the team in the long term.
## The Last Stage
At the end of the day, keep trying things out, have fun, and remember—it's called TigerBeetle, not
only because it's fast, but because it's small!
> You dont really suppose, do you, that all your adventures and escapes were managed by mere luck,
> just for your sole benefit? You are a very fine person, Mr. Baggins, and I am very fond of you;
> but you are only quite a little fellow in a wide world after all!”
>
> “Thank goodness!” said Bilbo laughing, and handed him the tobacco-jar.

View File

@@ -122,7 +122,11 @@ pub const Document = struct {
} }
/// Serialize the full document (length-prefixed) into `out`. /// Serialize the full document (length-prefixed) into `out`.
pub fn to_bytes(self: *const Document, gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8)) !void { pub fn to_bytes(
self: *const Document,
gpa: std.mem.Allocator,
out: *std.ArrayListUnmanaged(u8),
) !void {
try write_doc(self.pairs, gpa, out); try write_doc(self.pairs, gpa, out);
} }
}; };
@@ -157,7 +161,12 @@ const Parser = struct {
const ParseError = error{ InvalidBson, OutOfMemory }; const ParseError = error{ InvalidBson, OutOfMemory };
fn parse_doc_into(allocator: std.mem.Allocator, bytes: []const u8, idx: *usize, borrow: bool) ParseError![]const Pair { fn parse_doc_into(
allocator: std.mem.Allocator,
bytes: []const u8,
idx: *usize,
borrow: bool,
) ParseError![]const Pair {
const p = Parser{ .allocator = allocator, .bytes = bytes, .borrow = borrow }; const p = Parser{ .allocator = allocator, .bytes = bytes, .borrow = borrow };
return parse_doc_inner(p, idx); return parse_doc_inner(p, idx);
} }
@@ -356,7 +365,10 @@ fn parse_array(p: Parser, idx: *usize) ParseError![]const Value {
/// A borrowed parse of `bytes` into an arena: keys, strings, binary and /// A borrowed parse of `bytes` into an arena: keys, strings, binary and
/// regex payloads point into `bytes`; only the pair/value skeleton is /// regex payloads point into `bytes`; only the pair/value skeleton is
/// allocated. The result is valid while both `bytes` and `arena` live. /// allocated. The result is valid while both `bytes` and `arena` live.
pub fn spine(allocator: std.mem.Allocator, bytes: []const u8) error{ InvalidBson, OutOfMemory }![]const Pair { pub fn spine(
allocator: std.mem.Allocator,
bytes: []const u8,
) error{ InvalidBson, OutOfMemory }![]const Pair {
var idx: usize = 0; var idx: usize = 0;
return parse_doc_into(allocator, bytes, &idx, true); return parse_doc_into(allocator, bytes, &idx, true);
} }
@@ -414,7 +426,12 @@ pub fn skip_value(bytes: []const u8, idx: *usize, tag: u8) error{InvalidBson}!vo
/// Read one value of `tag` into a Value whose leaves borrow `bytes`; nested /// Read one value of `tag` into a Value whose leaves borrow `bytes`; nested
/// documents and arrays materialize their spines into `arena`. Advances /// documents and arrays materialize their spines into `arena`. Advances
/// `idx` past the value. /// `idx` past the value.
pub fn read_value(allocator: std.mem.Allocator, bytes: []const u8, idx: *usize, tag: u8) error{ InvalidBson, OutOfMemory }!Value { pub fn read_value(
allocator: std.mem.Allocator,
bytes: []const u8,
idx: *usize,
tag: u8,
) error{ InvalidBson, OutOfMemory }!Value {
switch (tag) { switch (tag) {
0x01 => { 0x01 => {
try ensure_available(bytes, idx.*, 8); try ensure_available(bytes, idx.*, 8);
@@ -542,7 +559,11 @@ pub fn read_value(allocator: std.mem.Allocator, bytes: []const u8, idx: *usize,
/// The value stored under `key` in a document's bytes, or null when absent. /// The value stored under `key` in a document's bytes, or null when absent.
/// Nested documents and arrays materialize their spines into `arena`. /// Nested documents and arrays materialize their spines into `arena`.
pub fn get_at(arena: std.mem.Allocator, bytes: []const u8, key: []const u8) error{ InvalidBson, OutOfMemory }!?Value { pub fn get_at(
arena: std.mem.Allocator,
bytes: []const u8,
key: []const u8,
) error{ InvalidBson, OutOfMemory }!?Value {
var idx: usize = 4; var idx: usize = 4;
while (idx + 1 < bytes.len and bytes[idx] != 0) { while (idx + 1 < bytes.len and bytes[idx] != 0) {
const tag = bytes[idx]; const tag = bytes[idx];
@@ -565,7 +586,11 @@ pub const SerializeError = error{
OutOfMemory, OutOfMemory,
}; };
pub fn write_value(v: Value, gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8)) SerializeError!void { pub fn write_value(
v: Value,
gpa: std.mem.Allocator,
out: *std.ArrayListUnmanaged(u8),
) SerializeError!void {
switch (v) { switch (v) {
.double => |d| { .double => |d| {
var buf: [8]u8 = undefined; var buf: [8]u8 = undefined;
@@ -618,13 +643,21 @@ pub fn write_value(v: Value, gpa: std.mem.Allocator, out: *std.ArrayListUnmanage
} }
} }
pub fn write_cstring(s: []const u8, gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8)) SerializeError!void { pub fn write_cstring(
s: []const u8,
gpa: std.mem.Allocator,
out: *std.ArrayListUnmanaged(u8),
) SerializeError!void {
if (std.mem.indexOfScalar(u8, s, 0) != null) return error.BsonNulInKey; if (std.mem.indexOfScalar(u8, s, 0) != null) return error.BsonNulInKey;
try out.appendSlice(gpa, s); try out.appendSlice(gpa, s);
try out.append(gpa, 0); try out.append(gpa, 0);
} }
pub fn write_string(s: []const u8, gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8)) SerializeError!void { pub fn write_string(
s: []const u8,
gpa: std.mem.Allocator,
out: *std.ArrayListUnmanaged(u8),
) SerializeError!void {
if (s.len + 1 > std.math.maxInt(u32)) return error.BsonTooLarge; if (s.len + 1 > std.math.maxInt(u32)) return error.BsonTooLarge;
var buf: [4]u8 = undefined; var buf: [4]u8 = undefined;
std.mem.writeInt(u32, &buf, @intCast(s.len + 1), .little); std.mem.writeInt(u32, &buf, @intCast(s.len + 1), .little);
@@ -633,7 +666,11 @@ pub fn write_string(s: []const u8, gpa: std.mem.Allocator, out: *std.ArrayListUn
try out.append(gpa, 0); try out.append(gpa, 0);
} }
pub fn write_element(pair: Pair, gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8)) SerializeError!void { pub fn write_element(
pair: Pair,
gpa: std.mem.Allocator,
out: *std.ArrayListUnmanaged(u8),
) SerializeError!void {
try out.append(gpa, pair.value.type_tag()); try out.append(gpa, pair.value.type_tag());
try write_cstring(pair.key, gpa, out); try write_cstring(pair.key, gpa, out);
try write_value(pair.value, gpa, out); try write_value(pair.value, gpa, out);
@@ -648,7 +685,11 @@ fn begin_frame(gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8)) Seriali
} }
/// Terminate the frame opened at `len_pos` and patch in its total length. /// Terminate the frame opened at `len_pos` and patch in its total length.
fn end_frame(gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8), len_pos: usize) SerializeError!void { fn end_frame(
gpa: std.mem.Allocator,
out: *std.ArrayListUnmanaged(u8),
len_pos: usize,
) SerializeError!void {
try out.append(gpa, 0); try out.append(gpa, 0);
const total = out.items.len - len_pos; const total = out.items.len - len_pos;
if (total > std.math.maxInt(u32)) return error.BsonTooLarge; if (total > std.math.maxInt(u32)) return error.BsonTooLarge;
@@ -656,13 +697,21 @@ fn end_frame(gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8), len_pos:
} }
/// Write a length-prefixed document. Length is patched in after the body. /// Write a length-prefixed document. Length is patched in after the body.
pub fn write_doc(pairs: []const Pair, gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8)) SerializeError!void { pub fn write_doc(
pairs: []const Pair,
gpa: std.mem.Allocator,
out: *std.ArrayListUnmanaged(u8),
) SerializeError!void {
const len_pos = try begin_frame(gpa, out); const len_pos = try begin_frame(gpa, out);
for (pairs) |p| try write_element(p, gpa, out); for (pairs) |p| try write_element(p, gpa, out);
try end_frame(gpa, out, len_pos); try end_frame(gpa, out, len_pos);
} }
fn write_array(items: []const Value, gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8)) SerializeError!void { fn write_array(
items: []const Value,
gpa: std.mem.Allocator,
out: *std.ArrayListUnmanaged(u8),
) SerializeError!void {
const len_pos = try begin_frame(gpa, out); const len_pos = try begin_frame(gpa, out);
var buf: [16]u8 = undefined; var buf: [16]u8 = undefined;
for (items, 0..) |item, i| { for (items, 0..) |item, i| {
@@ -686,7 +735,11 @@ pub fn serialize_value(gpa: std.mem.Allocator, v: Value) ![]u8 {
/// Append the serialized-key bytes of `v` (type tag + payload) to `out`. The /// Append the serialized-key bytes of `v` (type tag + payload) to `out`. The
/// appending form of serialize_value, for callers reusing one scratch buffer /// appending form of serialize_value, for callers reusing one scratch buffer
/// across many keys. /// across many keys.
pub fn write_serialized_value(v: Value, gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8)) !void { pub fn write_serialized_value(
v: Value,
gpa: std.mem.Allocator,
out: *std.ArrayListUnmanaged(u8),
) !void {
try out.append(gpa, v.type_tag()); try out.append(gpa, v.type_tag());
try write_value(v, gpa, out); try write_value(v, gpa, out);
} }
@@ -709,7 +762,10 @@ pub fn copy_value(arena: std.mem.Allocator, v: Value) std.mem.Allocator.Error!Va
}; };
} }
pub fn copy_pairs(arena: std.mem.Allocator, pairs: []const Pair) std.mem.Allocator.Error![]const Pair { pub fn copy_pairs(
arena: std.mem.Allocator,
pairs: []const Pair,
) std.mem.Allocator.Error![]const Pair {
const out = try arena.alloc(Pair, pairs.len); const out = try arena.alloc(Pair, pairs.len);
for (pairs, 0..) |p, i| { for (pairs, 0..) |p, i| {
out[i] = .{ .key = try arena.dupe(u8, p.key), .value = try copy_value(arena, p.value) }; out[i] = .{ .key = try arena.dupe(u8, p.key), .value = try copy_value(arena, p.value) };
@@ -717,7 +773,10 @@ pub fn copy_pairs(arena: std.mem.Allocator, pairs: []const Pair) std.mem.Allocat
return out; return out;
} }
fn copy_values(arena: std.mem.Allocator, items: []const Value) std.mem.Allocator.Error![]const Value { fn copy_values(
arena: std.mem.Allocator,
items: []const Value,
) std.mem.Allocator.Error![]const Value {
const out = try arena.alloc(Value, items.len); const out = try arena.alloc(Value, items.len);
for (items, 0..) |item, i| out[i] = try copy_value(arena, item); for (items, 0..) |item, i| out[i] = try copy_value(arena, item);
return out; return out;
@@ -879,7 +938,11 @@ pub fn encoded_leading_datetime(key: []const u8) ?i64 {
/// ambiguous. Escaping `00` as `00 FF` fixes both problems at once: a real /// ambiguous. Escaping `00` as `00 FF` fixes both problems at once: a real
/// NUL encodes above the `00 00` terminator, and any byte >= 01 is above it /// NUL encodes above the `00 00` terminator, and any byte >= 01 is above it
/// too, so "shorter is less" falls out to match `std.mem.order`. /// too, so "shorter is less" falls out to match `std.mem.order`.
fn encode_escaped(bytes: []const u8, gpa: std.mem.Allocator, out: *std.ArrayListUnmanaged(u8)) !void { fn encode_escaped(
bytes: []const u8,
gpa: std.mem.Allocator,
out: *std.ArrayListUnmanaged(u8),
) !void {
for (bytes) |b| { for (bytes) |b| {
try out.append(gpa, b); try out.append(gpa, b);
if (b == 0x00) try out.append(gpa, 0xFF); if (b == 0x00) try out.append(gpa, 0xFF);
@@ -1128,23 +1191,47 @@ test "encode_key order matches bson.compare on every pair" {
.min_key, .min_key,
.null, .null,
// Numbers: cross-type equality, sign, zero, extremes, NaN. // Numbers: cross-type equality, sign, zero, extremes, NaN.
.{ .int32 = -2147483648 }, .{ .int32 = -1 }, .{ .int32 = 0 }, .{ .int32 = -2147483648 },
.{ .int32 = 1 }, .{ .int32 = 2 }, .{ .int32 = 2147483647 }, .{ .int32 = -1 },
.{ .int64 = -9223372036854775807 }, .{ .int64 = -1 }, .{ .int64 = 0 }, .{ .int32 = 0 },
.{ .int64 = 1 }, .{ .int64 = 9223372036854775807 }, .{ .int32 = 1 },
.{ .double = -std.math.inf(f64) }, .{ .double = -1.5 }, .{ .double = -0.0 }, .{ .int32 = 2 },
.{ .double = 0.0 }, .{ .double = 0.5 }, .{ .double = 1.0 }, .{ .int32 = 2147483647 },
.{ .double = 1.5 }, .{ .double = std.math.inf(f64) }, .{ .int64 = -9223372036854775807 },
.{ .int64 = -1 },
.{ .int64 = 0 },
.{ .int64 = 1 },
.{ .int64 = 9223372036854775807 },
.{ .double = -std.math.inf(f64) },
.{ .double = -1.5 },
.{ .double = -0.0 },
.{ .double = 0.0 },
.{ .double = 0.5 },
.{ .double = 1.0 },
.{ .double = 1.5 },
.{ .double = std.math.inf(f64) },
.{ .double = std.math.nan(f64) }, .{ .double = std.math.nan(f64) },
// Strings, including embedded NUL and prefix relationships. // Strings, including embedded NUL and prefix relationships.
.{ .string = "" }, .{ .string = "\x00" }, .{ .string = "\x00b" }, .{ .string = "" },
.{ .string = "a" }, .{ .string = "a\x00" }, .{ .string = "a\x00b" }, .{ .string = "\x00" },
.{ .string = "ab" }, .{ .string = "b" }, .{ .string = "\xff" }, .{ .string = "\x00b" },
.{ .string = "a" },
.{ .string = "a\x00" },
.{ .string = "a\x00b" },
.{ .string = "ab" },
.{ .string = "b" },
.{ .string = "\xff" },
// Same rank as string, so these must interleave with them. // Same rank as string, so these must interleave with them.
.{ .symbol = "a" }, .{ .code = "ab" }, .{ .symbol = "a" },
.{ .doc = &.{} }, .{ .doc = &nested }, .{ .doc = &nested2 }, .{ .code = "ab" },
.{ .doc = &nested_l }, .{ .doc = &two_pairs }, .{ .doc = &.{} },
.{ .array = &.{} }, .{ .array = &arr1 }, .{ .array = &arr2 }, .{ .doc = &nested },
.{ .doc = &nested2 },
.{ .doc = &nested_l },
.{ .doc = &two_pairs },
.{ .array = &.{} },
.{ .array = &arr1 },
.{ .array = &arr2 },
.{ .array = &arr_str }, .{ .array = &arr_str },
// Binary orders by length first, then bytes, then subtype. // Binary orders by length first, then bytes, then subtype.
.{ .binary = .{ .subtype = 0, .data = "" } }, .{ .binary = .{ .subtype = 0, .data = "" } },
@@ -1155,11 +1242,15 @@ test "encode_key order matches bson.compare on every pair" {
.{ .object_id = [_]u8{0} ** 12 }, .{ .object_id = [_]u8{0} ** 12 },
.{ .object_id = [_]u8{0} ** 11 ++ [_]u8{1} }, .{ .object_id = [_]u8{0} ** 11 ++ [_]u8{1} },
.{ .object_id = [_]u8{255} ** 12 }, .{ .object_id = [_]u8{255} ** 12 },
.{ .bool = false }, .{ .bool = true }, .{ .bool = false },
.{ .datetime = std.math.minInt(i64) }, .{ .datetime = -1 }, .{ .bool = true },
.{ .datetime = 0 }, .{ .datetime = 1 }, .{ .datetime = std.math.minInt(i64) },
.{ .datetime = -1 },
.{ .datetime = 0 },
.{ .datetime = 1 },
.{ .datetime = std.math.maxInt(i64) }, .{ .datetime = std.math.maxInt(i64) },
.{ .timestamp = 0 }, .{ .timestamp = 1 }, .{ .timestamp = 0 },
.{ .timestamp = 1 },
.{ .timestamp = std.math.maxInt(u64) }, .{ .timestamp = std.math.maxInt(u64) },
.{ .regex = .{ .pattern = "a", .options = "" } }, .{ .regex = .{ .pattern = "a", .options = "" } },
.{ .regex = .{ .pattern = "a", .options = "i" } }, .{ .regex = .{ .pattern = "a", .options = "i" } },

View File

@@ -320,7 +320,13 @@ pub const Engine = struct {
/// shared); the collection lock is acquired before the exclusive catalog /// shared); the collection lock is acquired before the exclusive catalog
/// lock is dropped, so a concurrent drop can never free it underneath. /// lock is dropped, so a concurrent drop can never free it underneath.
/// Returns null when the collection does not exist (and create is off). /// Returns null when the collection does not exist (and create is off).
pub fn lock_collection(self: *Engine, db_name: []const u8, coll_name: []const u8, write: bool, create: bool) !?*Collection { pub fn lock_collection(
self: *Engine,
db_name: []const u8,
coll_name: []const u8,
write: bool,
create: bool,
) !?*Collection {
var coll = self.get_collection(db_name, coll_name); var coll = self.get_collection(db_name, coll_name);
if (coll == null and create) { if (coll == null and create) {
self.catalog_lock.unlockShared(self.io); self.catalog_lock.unlockShared(self.io);
@@ -425,7 +431,13 @@ pub const Engine = struct {
/// Log an append (and its seq increment) under the log lock, marking /// Log an append (and its seq increment) under the log lock, marking
/// the append as in flight so a commit leader's seal covers it. /// the append as in flight so a commit leader's seal covers it.
fn log_append(self: *Engine, comptime kind: LogKind, db: []const u8, coll: []const u8, doc: []const u8) !void { fn log_append(
self: *Engine,
comptime kind: LogKind,
db: []const u8,
coll: []const u8,
doc: []const u8,
) !void {
_ = self.pending_appends.fetchAdd(1, .acq_rel); _ = self.pending_appends.fetchAdd(1, .acq_rel);
defer { defer {
// The increment above pairs with this decrement on every return // The increment above pairs with this decrement on every return
@@ -468,12 +480,24 @@ pub const Engine = struct {
/// Insert a document. Fails with error.DuplicateKey if the _id exists. /// Insert a document. Fails with error.DuplicateKey if the _id exists.
/// Generates an ObjectId _id when absent. /// Generates an ObjectId _id when absent.
pub fn insert(self: *Engine, db_name: []const u8, coll_name: []const u8, doc: *const bson.Document, oid_gen: *bson.ObjectIdGen) !void { pub fn insert(
self: *Engine,
db_name: []const u8,
coll_name: []const u8,
doc: *const bson.Document,
oid_gen: *bson.ObjectIdGen,
) !void {
return self.upsert(db_name, coll_name, doc, oid_gen, .insert); return self.upsert(db_name, coll_name, doc, oid_gen, .insert);
} }
/// Insert or replace a document by _id (upsert without existence check). /// Insert or replace a document by _id (upsert without existence check).
pub fn replace(self: *Engine, db_name: []const u8, coll_name: []const u8, doc: *const bson.Document, oid_gen: *bson.ObjectIdGen) !void { pub fn replace(
self: *Engine,
db_name: []const u8,
coll_name: []const u8,
doc: *const bson.Document,
oid_gen: *bson.ObjectIdGen,
) !void {
return self.upsert(db_name, coll_name, doc, oid_gen, .replace); return self.upsert(db_name, coll_name, doc, oid_gen, .replace);
} }
@@ -574,7 +598,12 @@ pub const Engine = struct {
/// Remove a document by its `_id` value. Returns true if it existed. /// Remove a document by its `_id` value. Returns true if it existed.
/// The serialized-key encoding stays private to the engine. /// The serialized-key encoding stays private to the engine.
pub fn remove_by_id(self: *Engine, db_name: []const u8, coll_name: []const u8, id: bson.Value) !bool { pub fn remove_by_id(
self: *Engine,
db_name: []const u8,
coll_name: []const u8,
id: bson.Value,
) !bool {
const id_key = try bson.serialize_value(self.gpa, id); const id_key = try bson.serialize_value(self.gpa, id);
defer self.gpa.free(id_key); defer self.gpa.free(id_key);
return self.remove(db_name, coll_name, id_key); return self.remove(db_name, coll_name, id_key);
@@ -610,7 +639,12 @@ pub const Engine = struct {
return db.collections.get(coll_name); return db.collections.get(coll_name);
} }
pub fn get_doc(self: *Engine, db_name: []const u8, coll_name: []const u8, id_key: []const u8) ?[]const u8 { pub fn get_doc(
self: *Engine,
db_name: []const u8,
coll_name: []const u8,
id_key: []const u8,
) ?[]const u8 {
const coll = self.get_collection(db_name, coll_name) orelse return null; const coll = self.get_collection(db_name, coll_name) orelse return null;
const off = coll.docs.get(id_key) orelse return null; const off = coll.docs.get(id_key) orelse return null;
return coll.doc_bytes(off); return coll.doc_bytes(off);
@@ -636,7 +670,12 @@ pub const Engine = struct {
/// after the index builds over the existing documents and passes /// after the index builds over the existing documents and passes
/// uniqueness, so a rejected create persists nothing. Returns the new /// uniqueness, so a rejected create persists nothing. Returns the new
/// index (or the existing one when the spec matches — idempotent). /// index (or the existing one when the spec matches — idempotent).
pub fn create_index(self: *Engine, db_name: []const u8, coll_name: []const u8, spec_doc: *const bson.Document) !*index.Index { pub fn create_index(
self: *Engine,
db_name: []const u8,
coll_name: []const u8,
spec_doc: *const bson.Document,
) !*index.Index {
const coll = try self.get_or_create_collection(db_name, coll_name); const coll = try self.get_or_create_collection(db_name, coll_name);
var ix = try index.parse_spec(self.gpa, spec_doc); var ix = try index.parse_spec(self.gpa, spec_doc);
var committed = false; var committed = false;
@@ -675,7 +714,12 @@ pub const Engine = struct {
/// Remove a secondary index by name, persisting a drop record first. /// Remove a secondary index by name, persisting a drop record first.
/// Returns false when no such index exists. /// Returns false when no such index exists.
pub fn drop_index(self: *Engine, db_name: []const u8, coll_name: []const u8, index_name: []const u8) !bool { pub fn drop_index(
self: *Engine,
db_name: []const u8,
coll_name: []const u8,
index_name: []const u8,
) !bool {
const db = self.dbs.get(db_name) orelse return false; const db = self.dbs.get(db_name) orelse return false;
const coll = db.collections.get(coll_name) orelse return false; const coll = db.collections.get(coll_name) orelse return false;
if (coll.find_index(index_name) == null) return false; if (coll.find_index(index_name) == null) return false;
@@ -722,7 +766,13 @@ pub const Engine = struct {
/// Sweep one collection under its write lock; the lock is released on /// Sweep one collection under its write lock; the lock is released on
/// every return path. Returns how many documents were removed. /// every return path. Returns how many documents were removed.
fn ttl_sweep_coll(self: *Engine, coll: *Collection, now_ms: i64, db_name: []const u8, coll_name: []const u8) !usize { fn ttl_sweep_coll(
self: *Engine,
coll: *Collection,
now_ms: i64,
db_name: []const u8,
coll_name: []const u8,
) !usize {
try coll.lock.lock(self.io); try coll.lock.lock(self.io);
defer coll.lock.unlock(self.io); defer coll.lock.unlock(self.io);
// Ids are duped rather than aliased: `remove` frees the docs-map key // Ids are duped rather than aliased: `remove` frees the docs-map key
@@ -778,7 +828,11 @@ pub const Engine = struct {
while (it.next()) |entry| try out.append(self.gpa, entry.key_ptr.*); while (it.next()) |entry| try out.append(self.gpa, entry.key_ptr.*);
} }
pub fn collection_names(self: *Engine, db_name: []const u8, out: *std.ArrayListUnmanaged([]const u8)) !void { pub fn collection_names(
self: *Engine,
db_name: []const u8,
out: *std.ArrayListUnmanaged([]const u8),
) !void {
const db = self.dbs.get(db_name) orelse return; const db = self.dbs.get(db_name) orelse return;
var it = db.collections.iterator(); var it = db.collections.iterator();
while (it.next()) |entry| try out.append(self.gpa, entry.key_ptr.*); while (it.next()) |entry| try out.append(self.gpa, entry.key_ptr.*);
@@ -786,7 +840,11 @@ pub const Engine = struct {
// -- internals ----------------------------------------------------------- // -- internals -----------------------------------------------------------
pub fn get_or_create_collection(self: *Engine, db_name: []const u8, coll_name: []const u8) !*Collection { pub fn get_or_create_collection(
self: *Engine,
db_name: []const u8,
coll_name: []const u8,
) !*Collection {
const db = self.dbs.getPtr(db_name) orelse { const db = self.dbs.getPtr(db_name) orelse {
const db_key = try self.gpa.dupe(u8, db_name); const db_key = try self.gpa.dupe(u8, db_name);
errdefer self.gpa.free(db_key); errdefer self.gpa.free(db_key);
@@ -808,7 +866,11 @@ pub const Engine = struct {
/// generated ObjectId `_id` when absent. /// generated ObjectId `_id` when absent.
/// The canonical bytes of `doc`, with an ObjectId `_id` generated when /// The canonical bytes of `doc`, with an ObjectId `_id` generated when
/// absent. The result is owned by the caller. /// absent. The result is owned by the caller.
fn serialize_with_id(self: *Engine, doc: *const bson.Document, oid_gen: *bson.ObjectIdGen) ![]u8 { fn serialize_with_id(
self: *Engine,
doc: *const bson.Document,
oid_gen: *bson.ObjectIdGen,
) ![]u8 {
if (doc.get("_id") != null) return serialize_doc(self.gpa, doc); if (doc.get("_id") != null) return serialize_doc(self.gpa, doc);
var pairs: std.ArrayListUnmanaged(bson.Pair) = .empty; var pairs: std.ArrayListUnmanaged(bson.Pair) = .empty;
defer pairs.deinit(self.gpa); defer pairs.deinit(self.gpa);
@@ -994,7 +1056,13 @@ pub const Engine = struct {
/// Re-emit one collection's index specs and documents into the compacted /// Re-emit one collection's index specs and documents into the compacted
/// log, under the collection's write lock (released on every return /// log, under the collection's write lock (released on every return
/// path, including errors). /// path, including errors).
fn compact_snapshot_coll(self: *Engine, coll: *Collection, new_log: *storage.Log, db_name: []const u8, coll_name: []const u8) !void { fn compact_snapshot_coll(
self: *Engine,
coll: *Collection,
new_log: *storage.Log,
db_name: []const u8,
coll_name: []const u8,
) !void {
try coll.lock.lock(self.io); try coll.lock.lock(self.io);
defer coll.lock.unlock(self.io); defer coll.lock.unlock(self.io);
// Re-emit the index definitions first: a compacted log that dropped // Re-emit the index definitions first: a compacted log that dropped
@@ -1045,7 +1113,13 @@ pub const Engine = struct {
while (doc_it.next()) |doc_entry| { while (doc_it.next()) |doc_entry| {
ix.append_doc_entries(self.gpa, coll.doc_bytes(doc_entry.value_ptr.*), doc_entry.key_ptr.*) catch |err| switch (err) { ix.append_doc_entries(self.gpa, coll.doc_bytes(doc_entry.value_ptr.*), doc_entry.key_ptr.*) catch |err| switch (err) {
error.ParallelArrays => { error.ParallelArrays => {
std.debug.print("multiforadb: WARNING: index '{s}' cannot index an existing document; entry skipped\n", .{ix.name}); std.debug.print(
"multiforadb: WARNING: index '{s}' cannot index an existing " ++
"document; entry skipped\n",
.{
ix.name,
},
);
continue; continue;
}, },
else => return err, else => return err,
@@ -1053,13 +1127,23 @@ pub const Engine = struct {
} }
// Tolerated, not enforced: the database must always open. // Tolerated, not enforced: the database must always open.
if (try ix.finish_bulk(self.gpa, false)) { if (try ix.finish_bulk(self.gpa, false)) {
std.debug.print("multiforadb: WARNING: unique index '{s}' has duplicate keys in existing data; duplicates not enforced for existing documents\n", .{ix.name}); std.debug.print(
"multiforadb: WARNING: unique index '{s}' has duplicate keys in existing " ++
"data; duplicates not enforced for existing documents\n",
.{
ix.name,
},
);
} }
} }
/// Register an (empty) index from a persisted spec document. A repeated /// Register an (empty) index from a persisted spec document. A repeated
/// create record for the same name is an idempotent no-op. /// create record for the same name is an idempotent no-op.
fn register_index_from_spec(self: *Engine, coll: *Collection, spec_doc: *const bson.Document) !void { fn register_index_from_spec(
self: *Engine,
coll: *Collection,
spec_doc: *const bson.Document,
) !void {
var ix = try index.parse_spec(self.gpa, spec_doc); var ix = try index.parse_spec(self.gpa, spec_doc);
var committed = false; var committed = false;
defer if (!committed) ix.deinit(self.gpa); defer if (!committed) ix.deinit(self.gpa);
@@ -1103,7 +1187,9 @@ fn apply_record(ctx: *anyopaque, record: storage.Record, doc: *bson.Document) an
switch (record.type) { switch (record.type) {
storage.record_type_index_create => { storage.record_type_index_create => {
self.register_index_from_spec(coll, doc) catch |err| { self.register_index_from_spec(coll, doc) catch |err| {
std.debug.print("multiforadb: index create record failed to apply: {s}\n", .{@errorName(err)}); std.debug.print("multiforadb: index create record failed to apply: {s}\n", .{
@errorName(err),
});
return; return;
}; };
return; return;
@@ -1455,7 +1541,12 @@ test "concurrent readers and writers on a threaded Io" {
var remaining = std.atomic.Value(usize).init(@intCast(total)); var remaining = std.atomic.Value(usize).init(@intCast(total));
const Worker = struct { const Worker = struct {
fn writer(e: *Engine, id_counter: *std.atomic.Value(i32), pending: *std.atomic.Value(usize), alloc: std.mem.Allocator) error{Canceled}!void { fn writer(
e: *Engine,
id_counter: *std.atomic.Value(i32),
pending: *std.atomic.Value(usize),
alloc: std.mem.Allocator,
) error{Canceled}!void {
while (true) { while (true) {
const id = id_counter.fetchAdd(1, .monotonic); const id = id_counter.fetchAdd(1, .monotonic);
if (id > total) return; if (id > total) return;
@@ -1605,7 +1696,11 @@ test "concurrent writers compacting: the log survives a reopen" {
} }
} }
fn writer(e: *Engine, id_counter: *std.atomic.Value(i32), alloc: std.mem.Allocator) error{Canceled}!void { fn writer(
e: *Engine,
id_counter: *std.atomic.Value(i32),
alloc: std.mem.Allocator,
) error{Canceled}!void {
while (true) { while (true) {
const id = id_counter.fetchAdd(1, .monotonic); const id = id_counter.fetchAdd(1, .monotonic);
if (id > total) return; if (id > total) return;
@@ -1645,7 +1740,14 @@ test "concurrent writers compacting: the log survives a reopen" {
/// A spec document for a single-path index, built by serializing and /// A spec document for a single-path index, built by serializing and
/// re-parsing so the pairs are arena-owned. /// re-parsing so the pairs are arena-owned.
fn index_spec(gpa: std.mem.Allocator, path: []const u8, name: []const u8, unique: bool, sparse: bool, ttl: ?i64) !bson.Document { fn index_spec(
gpa: std.mem.Allocator,
path: []const u8,
name: []const u8,
unique: bool,
sparse: bool,
ttl: ?i64,
) !bson.Document {
var out: std.ArrayListUnmanaged(u8) = .empty; var out: std.ArrayListUnmanaged(u8) = .empty;
defer out.deinit(gpa); defer out.deinit(gpa);
var pairs: std.ArrayListUnmanaged(bson.Pair) = .empty; var pairs: std.ArrayListUnmanaged(bson.Pair) = .empty;
@@ -1662,7 +1764,14 @@ fn index_spec(gpa: std.mem.Allocator, path: []const u8, name: []const u8, unique
} }
/// Number of entries the named index has for a single-value equality key. /// Number of entries the named index has for a single-value equality key.
fn index_count(gpa: std.mem.Allocator, engine: *Engine, db_name: []const u8, coll_name: []const u8, name: []const u8, key_value: bson.Value) !usize { fn index_count(
gpa: std.mem.Allocator,
engine: *Engine,
db_name: []const u8,
coll_name: []const u8,
name: []const u8,
key_value: bson.Value,
) !usize {
const coll = engine.get_collection(db_name, coll_name) orelse return 0; const coll = engine.get_collection(db_name, coll_name) orelse return 0;
for (coll.indexes.items) |*ix| { for (coll.indexes.items) |*ix| {
if (std.mem.eql(u8, ix.name, name)) { if (std.mem.eql(u8, ix.name, name)) {

View File

@@ -113,7 +113,12 @@ fn run(seed: u64, ops: usize, max_len: usize) !void {
if (std.mem.eql(u8, got, d.id)) hit = true; if (std.mem.eql(u8, got, d.id)) hit = true;
} }
if (!hit) { if (!hit) {
std.debug.print("seed {d} op {d}: id {s} (key len {d}) not found by descent\n", .{ seed, op, d.id, d.s.len }); std.debug.print("seed {d} op {d}: id {s} (key len {d}) not found by descent\n", .{
seed,
op,
d.id,
d.s.len,
});
return error.EntryUnreachable; return error.EntryUnreachable;
} }
} }

View File

@@ -99,7 +99,10 @@ pub fn main(init: std.process.Init) !void {
var engine = try mongo.db.Engine.open(init.gpa, init.io, db_path); var engine = try mongo.db.Engine.open(init.gpa, init.io, db_path);
defer engine.deinit(); defer engine.deinit();
engine.compact_threshold = compact_threshold; engine.compact_threshold = compact_threshold;
std.debug.print("multiforadb: opened database '{s}' (compact threshold {d})\n", .{ db_path, compact_threshold }); std.debug.print("multiforadb: opened database '{s}' (compact threshold {d})\n", .{
db_path,
compact_threshold,
});
var server = mongo.server.Server{ var server = mongo.server.Server{
.gpa = init.gpa, .gpa = init.gpa,

View File

@@ -14,7 +14,11 @@ const bson = @import("bson.zig");
/// input — but it must still be a possible error, not a panic. /// input — but it must still be a possible error, not a panic.
pub const QueryError = error{ OutOfMemory, InvalidBson }; pub const QueryError = error{ OutOfMemory, InvalidBson };
pub fn matches(gpa: std.mem.Allocator, filter: *const bson.Document, doc: *const bson.Document) QueryError!bool { pub fn matches(
gpa: std.mem.Allocator,
filter: *const bson.Document,
doc: *const bson.Document,
) QueryError!bool {
for (filter.pairs) |p| { for (filter.pairs) |p| {
if (p.key.len > 0 and p.key[0] == '$') { if (p.key.len > 0 and p.key[0] == '$') {
if (!try match_top_level(gpa, p.key, p.value, doc)) return false; if (!try match_top_level(gpa, p.key, p.value, doc)) return false;
@@ -25,7 +29,12 @@ pub fn matches(gpa: std.mem.Allocator, filter: *const bson.Document, doc: *const
return true; return true;
} }
fn match_top_level(gpa: std.mem.Allocator, op: []const u8, value: bson.Value, doc: *const bson.Document) QueryError!bool { fn match_top_level(
gpa: std.mem.Allocator,
op: []const u8,
value: bson.Value,
doc: *const bson.Document,
) QueryError!bool {
if (std.mem.eql(u8, op, "$and") or std.mem.eql(u8, op, "$or")) { if (std.mem.eql(u8, op, "$and") or std.mem.eql(u8, op, "$or")) {
const want_and = std.mem.eql(u8, op, "$and"); const want_and = std.mem.eql(u8, op, "$and");
const filters = switch (value) { const filters = switch (value) {
@@ -89,7 +98,12 @@ fn is_operator_doc(value: bson.Value) ?[]const bson.Pair {
/// collection scan. /// collection scan.
const inline_candidates = 8; const inline_candidates = 8;
fn field_matches(gpa: std.mem.Allocator, path: []const u8, expected: bson.Value, doc: *const bson.Document) QueryError!bool { fn field_matches(
gpa: std.mem.Allocator,
path: []const u8,
expected: bson.Value,
doc: *const bson.Document,
) QueryError!bool {
var stack_fallback = std.heap.stackFallback(inline_candidates * @sizeOf(bson.Value), gpa); var stack_fallback = std.heap.stackFallback(inline_candidates * @sizeOf(bson.Value), gpa);
const alloc = stack_fallback.get(); const alloc = stack_fallback.get();
@@ -103,7 +117,12 @@ fn field_matches(gpa: std.mem.Allocator, path: []const u8, expected: bson.Value,
/// The byte counterpart of field_matches: collects values by walking the /// The byte counterpart of field_matches: collects values by walking the
/// canonical BSON element stream of a stored document, skipping by length /// canonical BSON element stream of a stored document, skipping by length
/// any field the filter does not name. /// any field the filter does not name.
fn field_matches_bytes(gpa: std.mem.Allocator, path: []const u8, expected: bson.Value, bytes: []const u8) QueryError!bool { fn field_matches_bytes(
gpa: std.mem.Allocator,
path: []const u8,
expected: bson.Value,
bytes: []const u8,
) QueryError!bool {
// An arena, not a stack fallback: the byte walker materializes nested // An arena, not a stack fallback: the byte walker materializes nested
// doc/array values (whole-array equality, embedded docs) into the // doc/array values (whole-array equality, embedded docs) into the
// allocator it is given, and those must be freed with it. // allocator it is given, and those must be freed with it.
@@ -120,7 +139,10 @@ fn field_matches_bytes(gpa: std.mem.Allocator, path: []const u8, expected: bson.
/// MongoDB applies queries to array elements as well as the array itself. /// MongoDB applies queries to array elements as well as the array itself.
/// Index the snapshot length, re-reading items each iteration: appending /// Index the snapshot length, re-reading items each iteration: appending
/// may reallocate the buffer, which would invalidate a captured slice. /// may reallocate the buffer, which would invalidate a captured slice.
fn expand_arrays(alloc: std.mem.Allocator, candidates: *std.ArrayListUnmanaged(bson.Value)) QueryError!void { fn expand_arrays(
alloc: std.mem.Allocator,
candidates: *std.ArrayListUnmanaged(bson.Value),
) QueryError!void {
const direct_count = candidates.items.len; const direct_count = candidates.items.len;
var i: usize = 0; var i: usize = 0;
while (i < direct_count) : (i += 1) { while (i < direct_count) : (i += 1) {
@@ -133,7 +155,11 @@ fn expand_arrays(alloc: std.mem.Allocator, candidates: *std.ArrayListUnmanaged(b
/// The operator/equality half of field matching, shared by the tree and /// The operator/equality half of field matching, shared by the tree and
/// byte collectors. /// byte collectors.
fn apply_expected(gpa: std.mem.Allocator, expected: bson.Value, candidates: []const bson.Value) QueryError!bool { fn apply_expected(
gpa: std.mem.Allocator,
expected: bson.Value,
candidates: []const bson.Value,
) QueryError!bool {
if (is_operator_doc(expected)) |pairs| { if (is_operator_doc(expected)) |pairs| {
// $options modifies $regex wherever it appears in the document, so // $options modifies $regex wherever it appears in the document, so
// it has to be known before any operator runs. // it has to be known before any operator runs.
@@ -166,7 +192,11 @@ fn apply_expected(gpa: std.mem.Allocator, expected: bson.Value, candidates: []co
/// byte-matcher counterpart of `matches`, used by scans. Same semantics, /// byte-matcher counterpart of `matches`, used by scans. Same semantics,
/// different collection: fields the filter does not name are skipped by /// different collection: fields the filter does not name are skipped by
/// length instead of materialized. /// length instead of materialized.
pub fn matches_bytes(gpa: std.mem.Allocator, filter: []const bson.Pair, bytes: []const u8) QueryError!bool { pub fn matches_bytes(
gpa: std.mem.Allocator,
filter: []const bson.Pair,
bytes: []const u8,
) QueryError!bool {
for (filter) |p| { for (filter) |p| {
if (p.key.len > 0 and p.key[0] == '$') { if (p.key.len > 0 and p.key[0] == '$') {
if (!try match_top_level_bytes(gpa, p.key, p.value, bytes)) return false; if (!try match_top_level_bytes(gpa, p.key, p.value, bytes)) return false;
@@ -177,7 +207,12 @@ pub fn matches_bytes(gpa: std.mem.Allocator, filter: []const bson.Pair, bytes: [
return true; return true;
} }
fn match_top_level_bytes(gpa: std.mem.Allocator, op: []const u8, value: bson.Value, bytes: []const u8) QueryError!bool { fn match_top_level_bytes(
gpa: std.mem.Allocator,
op: []const u8,
value: bson.Value,
bytes: []const u8,
) QueryError!bool {
if (std.mem.eql(u8, op, "$and") or std.mem.eql(u8, op, "$or")) { if (std.mem.eql(u8, op, "$and") or std.mem.eql(u8, op, "$or")) {
const want_and = std.mem.eql(u8, op, "$and"); const want_and = std.mem.eql(u8, op, "$and");
const filters = switch (value) { const filters = switch (value) {
@@ -215,7 +250,13 @@ fn match_top_level_bytes(gpa: std.mem.Allocator, op: []const u8, value: bson.Val
/// Collect values reachable at `path` from a document's canonical bytes — /// Collect values reachable at `path` from a document's canonical bytes —
/// the byte counterpart of `collect_values`, with the same traversal, the /// the byte counterpart of `collect_values`, with the same traversal, the
/// same order and the same multikey semantics. Appends into `out`. /// same order and the same multikey semantics. Appends into `out`.
pub fn collect_values_bytes(gpa: std.mem.Allocator, bytes: []const u8, path: []const u8, out: *std.ArrayListUnmanaged(bson.Value), depth: usize) QueryError!void { pub fn collect_values_bytes(
gpa: std.mem.Allocator,
bytes: []const u8,
path: []const u8,
out: *std.ArrayListUnmanaged(bson.Value),
depth: usize,
) QueryError!void {
var it = std.mem.splitScalar(u8, path, '.'); var it = std.mem.splitScalar(u8, path, '.');
const first = it.next() orelse return; const first = it.next() orelse return;
const rest = it.rest(); const rest = it.rest();
@@ -241,7 +282,15 @@ pub fn collect_values_bytes(gpa: std.mem.Allocator, bytes: []const u8, path: []c
} }
} }
fn collect_from_value_bytes(gpa: std.mem.Allocator, bytes: []const u8, idx: *usize, tag: u8, path: []const u8, out: *std.ArrayListUnmanaged(bson.Value), depth: usize) QueryError!void { fn collect_from_value_bytes(
gpa: std.mem.Allocator,
bytes: []const u8,
idx: *usize,
tag: u8,
path: []const u8,
out: *std.ArrayListUnmanaged(bson.Value),
depth: usize,
) QueryError!void {
if (depth > 8) { if (depth > 8) {
try bson.skip_value(bytes, idx, tag); try bson.skip_value(bytes, idx, tag);
return; return;
@@ -345,7 +394,13 @@ fn parse_op(name: []const u8) Op {
return op_names.get(name) orelse .unknown; return op_names.get(name) orelse .unknown;
} }
fn match_operator(gpa: std.mem.Allocator, op: Op, value: bson.Value, actuals: []const bson.Value, regex_options: []const u8) QueryError!bool { fn match_operator(
gpa: std.mem.Allocator,
op: Op,
value: bson.Value,
actuals: []const bson.Value,
regex_options: []const u8,
) QueryError!bool {
if (op == .eq) { if (op == .eq) {
for (actuals) |a| if (bson.compare(a, value) == .eq) return true; for (actuals) |a| if (bson.compare(a, value) == .eq) return true;
return false; return false;
@@ -489,7 +544,13 @@ fn match_operator(gpa: std.mem.Allocator, op: Op, value: bson.Value, actuals: []
/// Appends into `out`; on OOM, collection stops early (the engine is /// Appends into `out`; on OOM, collection stops early (the engine is
/// already failing at that point). Public because index entry generation /// already failing at that point). Public because index entry generation
/// must mirror field_matches exactly (src/index.zig). /// must mirror field_matches exactly (src/index.zig).
pub fn collect_values(gpa: std.mem.Allocator, pairs: []const bson.Pair, path: []const u8, out: *std.ArrayListUnmanaged(bson.Value), depth: usize) QueryError!void { pub fn collect_values(
gpa: std.mem.Allocator,
pairs: []const bson.Pair,
path: []const u8,
out: *std.ArrayListUnmanaged(bson.Value),
depth: usize,
) QueryError!void {
var it = std.mem.splitScalar(u8, path, '.'); var it = std.mem.splitScalar(u8, path, '.');
const first = it.next() orelse return; const first = it.next() orelse return;
@@ -506,7 +567,13 @@ pub fn collect_values(gpa: std.mem.Allocator, pairs: []const bson.Pair, path: []
} }
} }
fn collect_from_value(gpa: std.mem.Allocator, v: bson.Value, path: []const u8, out: *std.ArrayListUnmanaged(bson.Value), depth: usize) QueryError!void { fn collect_from_value(
gpa: std.mem.Allocator,
v: bson.Value,
path: []const u8,
out: *std.ArrayListUnmanaged(bson.Value),
depth: usize,
) QueryError!void {
if (depth > 8) return; if (depth > 8) return;
switch (v) { switch (v) {
.doc => |pairs| try collect_values(gpa, pairs, path, out, depth), .doc => |pairs| try collect_values(gpa, pairs, path, out, depth),
@@ -581,7 +648,14 @@ pub fn regex_match(pattern: []const u8, options: []const u8, text: []const u8) b
/// Match `pattern[p..]` against `text[t..]`, returning the new text /// Match `pattern[p..]` against `text[t..]`, returning the new text
/// position on success (null on failure). Backtracks via recursion. /// position on success (null on failure). Backtracks via recursion.
fn match_here(pattern: []const u8, p: *usize, text: []const u8, t: usize, ci: bool, dot_all: bool) ?usize { fn match_here(
pattern: []const u8,
p: *usize,
text: []const u8,
t: usize,
ci: bool,
dot_all: bool,
) ?usize {
var pos = t; var pos = t;
while (p.* < pattern.len) { while (p.* < pattern.len) {
const c = pattern[p.*]; const c = pattern[p.*];
@@ -669,7 +743,9 @@ fn match_here(pattern: []const u8, p: *usize, text: []const u8, t: usize, ci: bo
var q_end = element_end; var q_end = element_end;
var min: usize = 1; var min: usize = 1;
var max: usize = 1; var max: usize = 1;
if (element_end < pattern.len and (pattern[element_end] == '*' or pattern[element_end] == '+' or pattern[element_end] == '?')) { if (element_end < pattern.len and (pattern[element_end] == '*' or pattern[element_end] == '+' or pattern[
element_end
] == '?')) {
switch (pattern[element_end]) { switch (pattern[element_end]) {
'*' => { '*' => {
min = 0; min = 0;
@@ -847,7 +923,11 @@ const SortCtx = struct {
/// Pull each document's sort-key values into one flat allocation, so the /// Pull each document's sort-key values into one flat allocation, so the
/// comparator is pure and cannot fail. /// comparator is pure and cannot fail.
fn decorate(arena: std.mem.Allocator, docs: []*const bson.Document, keys: []const SortKey) QueryError![]SortedDoc { fn decorate(
arena: std.mem.Allocator,
docs: []*const bson.Document,
keys: []const SortKey,
) QueryError![]SortedDoc {
const entries = try arena.alloc(SortedDoc, docs.len); const entries = try arena.alloc(SortedDoc, docs.len);
const flat = try arena.alloc(bson.Value, docs.len * keys.len); const flat = try arena.alloc(bson.Value, docs.len * keys.len);
for (docs, 0..) |d, i| { for (docs, 0..) |d, i| {
@@ -861,7 +941,11 @@ fn decorate(arena: std.mem.Allocator, docs: []*const bson.Document, keys: []cons
} }
/// Sort `docs` in place by `keys`. /// Sort `docs` in place by `keys`.
pub fn sort_docs(arena: std.mem.Allocator, docs: []*const bson.Document, keys: []const SortKey) QueryError!void { pub fn sort_docs(
arena: std.mem.Allocator,
docs: []*const bson.Document,
keys: []const SortKey,
) QueryError!void {
if (keys.len == 0 or docs.len < 2) return; if (keys.len == 0 or docs.len < 2) return;
const entries = try decorate(arena, docs, keys); const entries = try decorate(arena, docs, keys);
std.mem.sort(SortedDoc, entries, SortCtx{ .keys = keys }, SortCtx.less); std.mem.sort(SortedDoc, entries, SortCtx{ .keys = keys }, SortCtx.less);
@@ -876,7 +960,12 @@ pub fn sort_docs(arena: std.mem.Allocator, docs: []*const bson.Document, keys: [
/// n log n comparisons to discard almost all of the result. This keeps a /// n log n comparisons to discard almost all of the result. This keeps a
/// k-element max-heap instead: one comparison against the heap root per /// k-element max-heap instead: one comparison against the heap root per
/// document, and only the survivors are ever ordered. /// document, and only the survivors are ever ordered.
pub fn sort_docs_top_k(arena: std.mem.Allocator, docs: []*const bson.Document, keys: []const SortKey, k: usize) QueryError!void { pub fn sort_docs_top_k(
arena: std.mem.Allocator,
docs: []*const bson.Document,
keys: []const SortKey,
k: usize,
) QueryError!void {
if (keys.len == 0 or docs.len < 2) return; if (keys.len == 0 or docs.len < 2) return;
if (k == 0) return; if (k == 0) return;
if (k >= docs.len) return sort_docs(arena, docs, keys); if (k >= docs.len) return sort_docs(arena, docs, keys);
@@ -936,7 +1025,12 @@ pub const ProjectionError = std.mem.Allocator.Error;
/// Apply a projection document, writing resulting pairs into `out` (which /// Apply a projection document, writing resulting pairs into `out` (which
/// should use the caller's arena so strings are owned). /// should use the caller's arena so strings are owned).
pub fn project(arena: std.mem.Allocator, doc: *const bson.Document, proj: *const bson.Document, out: *std.ArrayListUnmanaged(bson.Pair)) ProjectionError!void { pub fn project(
arena: std.mem.Allocator,
doc: *const bson.Document,
proj: *const bson.Document,
out: *std.ArrayListUnmanaged(bson.Pair),
) ProjectionError!void {
var inclusion: ?bool = null; var inclusion: ?bool = null;
var non_id_count: usize = 0; var non_id_count: usize = 0;
for (proj.pairs) |p| { for (proj.pairs) |p| {
@@ -984,7 +1078,12 @@ pub fn project(arena: std.mem.Allocator, doc: *const bson.Document, proj: *const
} }
/// Recursively apply exclusions to a nested document given the parent path. /// Recursively apply exclusions to a nested document given the parent path.
fn exclude_doc(arena: std.mem.Allocator, pairs: []const bson.Pair, proj: *const bson.Document, parent: []const u8) ProjectionError![]const bson.Pair { fn exclude_doc(
arena: std.mem.Allocator,
pairs: []const bson.Pair,
proj: *const bson.Document,
parent: []const u8,
) ProjectionError![]const bson.Pair {
var out: std.ArrayListUnmanaged(bson.Pair) = .empty; var out: std.ArrayListUnmanaged(bson.Pair) = .empty;
errdefer out.deinit(arena); errdefer out.deinit(arena);
for (pairs) |p| { for (pairs) |p| {
@@ -1024,7 +1123,12 @@ pub fn truthy(v: bson.Value) bool {
} }
/// Include a dotted path (e.g. "a.b.c"), creating nested documents as needed. /// Include a dotted path (e.g. "a.b.c"), creating nested documents as needed.
fn project_path(arena: std.mem.Allocator, pairs: []const bson.Pair, path: []const u8, out: *std.ArrayListUnmanaged(bson.Pair)) ProjectionError!void { fn project_path(
arena: std.mem.Allocator,
pairs: []const bson.Pair,
path: []const u8,
out: *std.ArrayListUnmanaged(bson.Pair),
) ProjectionError!void {
var it = std.mem.splitScalar(u8, path, '.'); var it = std.mem.splitScalar(u8, path, '.');
const first = it.next() orelse return; const first = it.next() orelse return;
const rest = it.rest(); const rest = it.rest();
@@ -1296,8 +1400,8 @@ test "first_value_at agrees with collect_values on its first element" {
const paths = [_][]const u8{ const paths = [_][]const u8{
"n", "sub", "sub.x", "sub.y", "sub.missing", "n", "sub", "sub.x", "sub.y", "sub.missing",
"items", "items.v", "items.0", "items.1.v", "items.9", "items", "items.v", "items.0", "items.1.v", "items.9",
"nums", "nums.0", "nums.1", "nums.5", "nums", "nums.0", "nums.1", "nums.5", "dup",
"dup", "missing", "n.deeper", "", "sub.x.y", "missing", "n.deeper", "", "sub.x.y",
}; };
for (paths) |path| { for (paths) |path| {
@@ -1308,12 +1412,17 @@ test "first_value_at agrees with collect_values on its first element" {
const first = first_value_at(d.pairs, path, 0); const first = first_value_at(d.pairs, path, 0);
if (list.items.len == 0) { if (list.items.len == 0) {
testing.expect(first == null) catch |e| { testing.expect(first == null) catch |e| {
std.debug.print("path '{s}': collect empty but first_value_at returned a value\n", .{path}); std.debug.print("path '{s}': collect empty but first_value_at returned a value\n", .{
path,
});
return e; return e;
}; };
} else { } else {
testing.expect(first != null) catch |e| { testing.expect(first != null) catch |e| {
std.debug.print("path '{s}': collect got {d} values but first_value_at returned null\n", .{ path, list.items.len }); std.debug.print("path '{s}': collect got {d} values but first_value_at returned null\n", .{
path,
list.items.len,
});
return e; return e;
}; };
testing.expectEqual(std.math.Order.eq, bson.compare(list.items[0], first.?)) catch |e| { testing.expectEqual(std.math.Order.eq, bson.compare(list.items[0], first.?)) catch |e| {
@@ -1372,7 +1481,11 @@ test "top-k selection matches a full sort on the leading page" {
const want = first_value_at(full[i].pairs, sk.path, 0) orelse bson.Value.null; const want = first_value_at(full[i].pairs, sk.path, 0) orelse bson.Value.null;
const got = first_value_at(topk[i].pairs, sk.path, 0) orelse bson.Value.null; const got = first_value_at(topk[i].pairs, sk.path, 0) orelse bson.Value.null;
testing.expectEqual(std.math.Order.eq, bson.compare(want, got)) catch |e| { testing.expectEqual(std.math.Order.eq, bson.compare(want, got)) catch |e| {
std.debug.print("k={d} pos={d} key='{s}' diverged from the full sort\n", .{ k, i, sk.path }); std.debug.print("k={d} pos={d} key='{s}' diverged from the full sort\n", .{
k,
i,
sk.path,
});
return e; return e;
}; };
} }
@@ -1482,7 +1595,7 @@ test "byte matcher agrees with the tree matcher on a corpus" {
np += 1; np += 1;
} }
if (rand.boolean()) { if (rand.boolean()) {
pairs[np] = .{ .key = "d", .value = .{ .doc = &.{ .{ .key = "e", .value = .{ .int32 = a } } } } }; pairs[np] = .{ .key = "d", .value = .{ .doc = &.{.{ .key = "e", .value = .{ .int32 = a } }} } };
np += 1; np += 1;
} }
var out: std.ArrayListUnmanaged(u8) = .empty; var out: std.ArrayListUnmanaged(u8) = .empty;
@@ -1541,7 +1654,11 @@ test "byte matcher agrees with the tree matcher on a corpus" {
const tree = try matches(gpa, &filter_doc, &doc); const tree = try matches(gpa, &filter_doc, &doc);
const byt = try matches_bytes(gpa, f_pairs.items, bytes); const byt = try matches_bytes(gpa, f_pairs.items, bytes);
if (tree != byt) { if (tree != byt) {
std.debug.print("case {d}: filter mismatch: tree={} bytes={}\n", .{ case, tree, byt }); std.debug.print("case {d}: filter mismatch: tree={} bytes={}\n", .{
case,
tree,
byt,
});
return error.ByteMatcherMismatch; return error.ByteMatcherMismatch;
} }
} }
@@ -1549,7 +1666,12 @@ test "byte matcher agrees with the tree matcher on a corpus" {
} }
/// Public single-value operator matcher, used by $pull and $elemMatch. /// Public single-value operator matcher, used by $pull and $elemMatch.
pub fn value_matches_operator(gpa: std.mem.Allocator, op: []const u8, value: bson.Value, actual: bson.Value) QueryError!bool { pub fn value_matches_operator(
gpa: std.mem.Allocator,
op: []const u8,
value: bson.Value,
actual: bson.Value,
) QueryError!bool {
var single: [1]bson.Value = .{actual}; var single: [1]bson.Value = .{actual};
return match_operator(gpa, parse_op(op), value, single[0..], ""); return match_operator(gpa, parse_op(op), value, single[0..], "");
} }

View File

@@ -120,7 +120,10 @@ fn handle_connection_inner(io: std.Io, stream: std.Io.net.Stream, server: *Serve
reader.interface.readSliceAll(&len_bytes) catch return; // clean client disconnect (EOF or RST) reader.interface.readSliceAll(&len_bytes) catch return; // clean client disconnect (EOF or RST)
const total: u32 = std.mem.readInt(u32, &len_bytes, .little); const total: u32 = std.mem.readInt(u32, &len_bytes, .little);
if (total < 16 or total > wire.max_message_size) { if (total < 16 or total > wire.max_message_size) {
std.debug.print("multiforadb: bad message length {d} on conn {d}\n", .{ total, connection_id }); std.debug.print("multiforadb: bad message length {d} on conn {d}\n", .{
total,
connection_id,
});
return; return;
} }
@@ -129,14 +132,22 @@ fn handle_connection_inner(io: std.Io, stream: std.Io.net.Stream, server: *Serve
msg_buf.items.len = total; msg_buf.items.len = total;
std.mem.writeInt(u32, msg_buf.items[0..4], total, .little); std.mem.writeInt(u32, msg_buf.items[0..4], total, .little);
reader.interface.readSliceAll(msg_buf.items[4..]) catch |err| { reader.interface.readSliceAll(msg_buf.items[4..]) catch |err| {
std.debug.print("multiforadb: read error on conn {d}: {s} (body, len {d})\n", .{ connection_id, @errorName(err), total }); std.debug.print("multiforadb: read error on conn {d}: {s} (body, len {d})\n", .{
connection_id,
@errorName(err),
total,
});
return; return;
}; };
var msg = wire.Message.parse(server.gpa, msg_buf.items) catch |err| { var msg = wire.Message.parse(server.gpa, msg_buf.items) catch |err| {
// Unparseable request: close the connection. // Unparseable request: close the connection.
const op: i32 = if (msg_buf.items.len >= 16) std.mem.readInt(i32, msg_buf.items[12..16], .little) else 0; const op: i32 = if (msg_buf.items.len >= 16) std.mem.readInt(i32, msg_buf.items[12..16], .little) else 0;
std.debug.print("multiforadb: bad message on conn {d}: {s} (opCode {d})\n", .{ connection_id, @errorName(err), op }); std.debug.print("multiforadb: bad message on conn {d}: {s} (opCode {d})\n", .{
connection_id,
@errorName(err),
op,
});
return; return;
}; };
defer msg.deinit(); defer msg.deinit();
@@ -146,7 +157,11 @@ fn handle_connection_inner(io: std.Io, stream: std.Io.net.Stream, server: *Serve
commands.dispatch(&ctx, &msg, &reply) catch |err| { commands.dispatch(&ctx, &msg, &reply) catch |err| {
// Discard any partial reply (the client would read the first // Discard any partial reply (the client would read the first
// ok field, which may already say 1) and send a clean error. // ok field, which may already say 1) and send a clean error.
std.debug.print("multiforadb: dispatch error on conn {d} cmd {s}: {s}\n", .{ connection_id, msg.command_name(), @errorName(err) }); std.debug.print("multiforadb: dispatch error on conn {d} cmd {s}: {s}\n", .{
connection_id,
msg.command_name(),
@errorName(err),
});
reply.pairs.clearRetainingCapacity(); reply.pairs.clearRetainingCapacity();
reply.put_error( reply.put_error(
@intFromEnum(commands.ErrorCode.internal_error), @intFromEnum(commands.ErrorCode.internal_error),
@@ -161,16 +176,25 @@ fn handle_connection_inner(io: std.Io, stream: std.Io.net.Stream, server: *Serve
else else
reply.build(server.gpa, reply_request_id, msg.request_id, &out_buf); reply.build(server.gpa, reply_request_id, msg.request_id, &out_buf);
built catch |err| { built catch |err| {
std.debug.print("multiforadb: reply build error on conn {d}: {s}\n", .{ connection_id, @errorName(err) }); std.debug.print("multiforadb: reply build error on conn {d}: {s}\n", .{
connection_id,
@errorName(err),
});
return; return;
}; };
reply_request_id +%= 1; reply_request_id +%= 1;
writer.interface.writeAll(out_buf.items) catch |err| { writer.interface.writeAll(out_buf.items) catch |err| {
std.debug.print("multiforadb: write error on conn {d}: {s}\n", .{ connection_id, @errorName(err) }); std.debug.print("multiforadb: write error on conn {d}: {s}\n", .{
connection_id,
@errorName(err),
});
return; return;
}; };
writer.interface.flush() catch |err| { writer.interface.flush() catch |err| {
std.debug.print("multiforadb: flush error on conn {d}: {s}\n", .{ connection_id, @errorName(err) }); std.debug.print("multiforadb: flush error on conn {d}: {s}\n", .{
connection_id,
@errorName(err),
});
return; return;
}; };
} }

View File

@@ -37,7 +37,12 @@ pub fn main() !void {
docs[i] = try doc_of(gpa, &pairs); docs[i] = try doc_of(gpa, &pairs);
_ = try ix.add_doc(gpa, docs[i], &[_]u8{ 'i', 'd', @intCast(i + 1) }, true); _ = try ix.add_doc(gpa, docs[i], &[_]u8{ 'i', 'd', @intCast(i + 1) }, true);
} }
std.debug.print("count={d} leaves={d} depth={d} overflow={d}\n", .{ ix.count(), ix.leaf_count, ix.depth, ix.overflow.items.len }); std.debug.print("count={d} leaves={d} depth={d} overflow={d}\n", .{
ix.count(),
ix.leaf_count,
ix.depth,
ix.overflow.items.len,
});
if (ix.overflow.items.len < 100_000) return error.NoSpill; if (ix.overflow.items.len < 100_000) return error.NoSpill;
// Every entry is found by exact key. // Every entry is found by exact key.

View File

@@ -48,7 +48,12 @@ pub fn main() !void {
const d = try doc_of(gpa, &pairs); const d = try doc_of(gpa, &pairs);
_ = try ix.add_doc(gpa, d, id, false); _ = try ix.add_doc(gpa, d, id, false);
} }
std.debug.print("count={d} leaves={d} depth={d} overflow={d}\n", .{ ix.count(), ix.leaf_count, ix.depth, ix.overflow.items.len }); std.debug.print("count={d} leaves={d} depth={d} overflow={d}\n", .{
ix.count(),
ix.leaf_count,
ix.depth,
ix.overflow.items.len,
});
if (ix.count() != N) return error.Bad; if (ix.count() != N) return error.Bad;
// Spot-check exact lookups. // Spot-check exact lookups.

View File

@@ -227,7 +227,10 @@ pub const Log = struct {
while (true) { while (true) {
var hdr: [block_header_len]u8 = undefined; var hdr: [block_header_len]u8 = undefined;
const n = self.file.readPositionalAll(self.io, &hdr, pos) catch |err| { const n = self.file.readPositionalAll(self.io, &hdr, pos) catch |err| {
std.debug.print("multiforadb: log read error at {d}: {s}\n", .{ pos, @errorName(err) }); std.debug.print("multiforadb: log read error at {d}: {s}\n", .{
pos,
@errorName(err),
});
return error.InvalidLog; return error.InvalidLog;
}; };
if (n == 0) return; // clean end if (n == 0) return; // clean end
@@ -262,7 +265,10 @@ pub const Log = struct {
codec_raw => try decomp.appendSlice(self.gpa, payload), codec_raw => try decomp.appendSlice(self.gpa, payload),
codec_lz4 => try lz4_decompress(self.gpa, payload, &decomp), codec_lz4 => try lz4_decompress(self.gpa, payload, &decomp),
else => { else => {
std.debug.print("multiforadb: unknown block codec {d} at {d}\n", .{ codec, pos }); std.debug.print("multiforadb: unknown block codec {d} at {d}\n", .{
codec,
pos,
});
return error.InvalidLog; return error.InvalidLog;
}, },
} }
@@ -280,7 +286,13 @@ pub const Log = struct {
/// and deliver it. Returns the record's byte length. Any framing failure /// and deliver it. Returns the record's byte length. Any framing failure
/// here is interior corruption: a block was sealed only with complete /// here is interior corruption: a block was sealed only with complete
/// records, and its hash proved the stored bytes intact. /// records, and its hash proved the stored bytes intact.
fn parse_record(self: *Log, bytes: []const u8, pos: u64, ctx: *anyopaque, callback: ReplayFn) !usize { fn parse_record(
self: *Log,
bytes: []const u8,
pos: u64,
ctx: *anyopaque,
callback: ReplayFn,
) !usize {
if (bytes.len < 4) return error.InvalidLog; if (bytes.len < 4) return error.InvalidLog;
const total: u32 = std.mem.readInt(u32, bytes[0..4], .little); const total: u32 = std.mem.readInt(u32, bytes[0..4], .little);
if (total < header_len) { if (total < header_len) {
@@ -321,26 +333,57 @@ pub const Log = struct {
return total; return total;
} }
pub fn append_upsert(self: *Log, db: []const u8, coll: []const u8, doc: []const u8, seq: u64) !void { pub fn append_upsert(
self: *Log,
db: []const u8,
coll: []const u8,
doc: []const u8,
seq: u64,
) !void {
try self.append(record_type_upsert, db, coll, doc, seq); try self.append(record_type_upsert, db, coll, doc, seq);
} }
pub fn append_delete(self: *Log, db: []const u8, coll: []const u8, doc: []const u8, seq: u64) !void { pub fn append_delete(
self: *Log,
db: []const u8,
coll: []const u8,
doc: []const u8,
seq: u64,
) !void {
try self.append(record_type_delete, db, coll, doc, seq); try self.append(record_type_delete, db, coll, doc, seq);
} }
/// The payload is the canonical index spec document ({v, key, name, /// The payload is the canonical index spec document ({v, key, name,
/// unique?, sparse?}); only apply_record interprets it. /// unique?, sparse?}); only apply_record interprets it.
pub fn append_index_create(self: *Log, db: []const u8, coll: []const u8, doc: []const u8, seq: u64) !void { pub fn append_index_create(
self: *Log,
db: []const u8,
coll: []const u8,
doc: []const u8,
seq: u64,
) !void {
try self.append(record_type_index_create, db, coll, doc, seq); try self.append(record_type_index_create, db, coll, doc, seq);
} }
/// The payload is {name: "..."}; only apply_record interprets it. /// The payload is {name: "..."}; only apply_record interprets it.
pub fn append_index_drop(self: *Log, db: []const u8, coll: []const u8, doc: []const u8, seq: u64) !void { pub fn append_index_drop(
self: *Log,
db: []const u8,
coll: []const u8,
doc: []const u8,
seq: u64,
) !void {
try self.append(record_type_index_drop, db, coll, doc, seq); try self.append(record_type_index_drop, db, coll, doc, seq);
} }
fn append(self: *Log, rtype: u8, db: []const u8, coll: []const u8, doc: []const u8, seq: u64) !void { fn append(
self: *Log,
rtype: u8,
db: []const u8,
coll: []const u8,
doc: []const u8,
seq: u64,
) !void {
if (std.mem.indexOfScalar(u8, db, 0) != null or std.mem.indexOfScalar(u8, coll, 0) != null) { if (std.mem.indexOfScalar(u8, db, 0) != null or std.mem.indexOfScalar(u8, coll, 0) != null) {
return error.NulInName; return error.NulInName;
} }
@@ -547,7 +590,11 @@ fn emit_literals(dst: []u8, literals: []const u8) usize {
/// Decompress an LZ4 block into `out` (appended). The block hash has /// Decompress an LZ4 block into `out` (appended). The block hash has
/// already proven the input intact when this runs during replay, so /// already proven the input intact when this runs during replay, so
/// structural failures here mean a bug or a raw-codec mismatch. /// structural failures here mean a bug or a raw-codec mismatch.
fn lz4_decompress(gpa: std.mem.Allocator, src: []const u8, out: *std.ArrayListUnmanaged(u8)) error{ CorruptLz4, OutOfMemory }!void { fn lz4_decompress(
gpa: std.mem.Allocator,
src: []const u8,
out: *std.ArrayListUnmanaged(u8),
) error{ CorruptLz4, OutOfMemory }!void {
var ip: usize = 0; var ip: usize = 0;
while (ip < src.len) { while (ip < src.len) {
const token = src[ip]; const token = src[ip];

View File

@@ -16,7 +16,15 @@ fn doc_of(gpa: std.mem.Allocator, pairs: []const bson.Pair) ![]u8 {
const Fact = struct { a: i32, b: i32 }; const Fact = struct { a: i32, b: i32 };
fn check_range(gpa: std.mem.Allocator, ix: *const index.Index, prefix: bson.Value, lo: ?bson.Value, hi: ?bson.Value, facts: []const Fact, alive: []const bool) !void { fn check_range(
gpa: std.mem.Allocator,
ix: *const index.Index,
prefix: bson.Value,
lo: ?bson.Value,
hi: ?bson.Value,
facts: []const Fact,
alive: []const bool,
) !void {
var out: std.ArrayListUnmanaged([]const u8) = .empty; var out: std.ArrayListUnmanaged([]const u8) = .empty;
defer out.deinit(gpa); defer out.deinit(gpa);
try ix.lookup_range(gpa, &.{prefix}, lo, true, hi, false, &out); try ix.lookup_range(gpa, &.{prefix}, lo, true, hi, false, &out);
@@ -29,7 +37,13 @@ fn check_range(gpa: std.mem.Allocator, ix: *const index.Index, prefix: bson.Valu
expected += 1; expected += 1;
} }
if (out.items.len != expected) { if (out.items.len != expected) {
std.debug.print("MISMATCH: prefix={d} lo={?d} hi={?d}: got {d}, want {d}\n", .{ prefix.int32, if (lo) |l| l.int32 else null, if (hi) |h| h.int32 else null, out.items.len, expected }); std.debug.print("MISMATCH: prefix={d} lo={?d} hi={?d}: got {d}, want {d}\n", .{
prefix.int32,
if (lo) |l| l.int32 else null,
if (hi) |h| h.int32 else null,
out.items.len,
expected,
});
std.process.exit(1); std.process.exit(1);
} }
} }
@@ -61,7 +75,12 @@ pub fn main() !void {
try ix.append_doc_entries(gpa, d, id); try ix.append_doc_entries(gpa, d, id);
} }
_ = try ix.finish_bulk(gpa, false); _ = try ix.finish_bulk(gpa, false);
std.debug.print("bulk: count={d} leaves={d} depth={d} nodes={d}\n", .{ ix.count(), ix.leaf_count, ix.depth, ix.nodes.items.len }); std.debug.print("bulk: count={d} leaves={d} depth={d} nodes={d}\n", .{
ix.count(),
ix.leaf_count,
ix.depth,
ix.nodes.items.len,
});
if (ix.count() != N) return error.BadCount; if (ix.count() != N) return error.BadCount;
// Random range checks against brute force. // Random range checks against brute force.
@@ -88,7 +107,12 @@ pub fn main() !void {
const d = try doc_of(gpa, &pairs); const d = try doc_of(gpa, &pairs);
_ = try ix.add_doc(gpa, d, id, false); _ = try ix.add_doc(gpa, d, id, false);
} }
std.debug.print("after inserts: count={d} leaves={d} depth={d} nodes={d}\n", .{ ix.count(), ix.leaf_count, ix.depth, ix.nodes.items.len }); std.debug.print("after inserts: count={d} leaves={d} depth={d} nodes={d}\n", .{
ix.count(),
ix.leaf_count,
ix.depth,
ix.nodes.items.len,
});
if (ix.count() != N + M) return error.BadCount; if (ix.count() != N + M) return error.BadCount;
for (0..500) |_| { for (0..500) |_| {
const a = rand.intRangeAtMost(i32, 0, 99); const a = rand.intRangeAtMost(i32, 0, 99);
@@ -113,7 +137,12 @@ pub fn main() !void {
ix.remove_doc(gpa, d, ids.items[i]); ix.remove_doc(gpa, d, ids.items[i]);
alive.items[i] = false; alive.items[i] = false;
} }
std.debug.print("after deletes: count={d} leaves={d} depth={d} nodes={d}\n", .{ ix.count(), ix.leaf_count, ix.depth, ix.nodes.items.len }); std.debug.print("after deletes: count={d} leaves={d} depth={d} nodes={d}\n", .{
ix.count(),
ix.leaf_count,
ix.depth,
ix.nodes.items.len,
});
if (ix.count() != (N + M) - (N + M) / 3 - 1) return error.BadCount; if (ix.count() != (N + M) - (N + M) / 3 - 1) return error.BadCount;
for (0..500) |_| { for (0..500) |_| {
const a = rand.intRangeAtMost(i32, 0, 99); const a = rand.intRangeAtMost(i32, 0, 99);
@@ -131,9 +160,15 @@ pub fn main() !void {
defer out.deinit(gpa); defer out.deinit(gpa);
try ix.lookup_eq(gpa, &.{.{ .int32 = a }}, &out); try ix.lookup_eq(gpa, &.{.{ .int32 = a }}, &out);
var expected: usize = 0; var expected: usize = 0;
for (facts.items, 0..) |f, fi| { if (alive.items[fi] and f.a == a) expected += 1; } for (facts.items, 0..) |f, fi| {
if (alive.items[fi] and f.a == a) expected += 1;
}
if (out.items.len != expected) { if (out.items.len != expected) {
std.debug.print("EQ MISMATCH a={d}: got {d} want {d}\n", .{ a, out.items.len, expected }); std.debug.print("EQ MISMATCH a={d}: got {d} want {d}\n", .{
a,
out.items.len,
expected,
});
return error.BadCount; return error.BadCount;
} }
} }
@@ -154,11 +189,19 @@ pub fn main() !void {
ix.remove_doc(gpa, d, ids.items[i]); ix.remove_doc(gpa, d, ids.items[i]);
remaining -= 1; remaining -= 1;
if (ix.count() != remaining) { if (ix.count() != remaining) {
std.debug.print("COUNT MISMATCH during drain: {d} != {d}\n", .{ ix.count(), remaining }); std.debug.print("COUNT MISMATCH during drain: {d} != {d}\n", .{
ix.count(),
remaining,
});
return error.BadCount; return error.BadCount;
} }
} }
std.debug.print("after full drain: count={d} leaves={d} depth={d} nodes={d}\n", .{ ix.count(), ix.leaf_count, ix.depth, ix.nodes.items.len }); std.debug.print("after full drain: count={d} leaves={d} depth={d} nodes={d}\n", .{
ix.count(),
ix.leaf_count,
ix.depth,
ix.nodes.items.len,
});
if (ix.count() != 0) return error.BadCount; if (ix.count() != 0) return error.BadCount;
// The drained tree still accepts and finds entries. // The drained tree still accepts and finds entries.
pairs[0] = .{ .key = "a", .value = .{ .int32 = 7 } }; pairs[0] = .{ .key = "a", .value = .{ .int32 = 7 } };

View File

@@ -21,7 +21,12 @@ pub fn apply(doc: *bson.Document, update: *const bson.Document) UpdateError!void
doc.pairs = try pairs.toOwnedSlice(arena); doc.pairs = try pairs.toOwnedSlice(arena);
} }
fn apply_operator(arena: std.mem.Allocator, pairs: *std.ArrayListUnmanaged(bson.Pair), op: []const u8, value: bson.Value) UpdateError!void { fn apply_operator(
arena: std.mem.Allocator,
pairs: *std.ArrayListUnmanaged(bson.Pair),
op: []const u8,
value: bson.Value,
) UpdateError!void {
if (std.mem.eql(u8, op, "$set")) { if (std.mem.eql(u8, op, "$set")) {
const ops = doc_pairs(value) orelse return error.InvalidUpdate; const ops = doc_pairs(value) orelse return error.InvalidUpdate;
for (ops) |p| { for (ops) |p| {
@@ -146,7 +151,11 @@ fn parse_index(seg: []const u8) ?usize {
} }
/// Shallow-copy a slice into a growable list backed by `arena`. /// Shallow-copy a slice into a growable list backed by `arena`.
fn copy_to_list(comptime T: type, arena: std.mem.Allocator, items: []const T) UpdateError!std.ArrayListUnmanaged(T) { fn copy_to_list(
comptime T: type,
arena: std.mem.Allocator,
items: []const T,
) UpdateError!std.ArrayListUnmanaged(T) {
var out: std.ArrayListUnmanaged(T) = .empty; var out: std.ArrayListUnmanaged(T) = .empty;
errdefer out.deinit(arena); errdefer out.deinit(arena);
try out.appendSlice(arena, items); try out.appendSlice(arena, items);
@@ -173,7 +182,12 @@ fn get_value(pairs: []const bson.Pair, segs: []const []const u8) ?bson.Value {
}; };
} }
fn set_path(arena: std.mem.Allocator, pairs: *std.ArrayListUnmanaged(bson.Pair), segs: []const []const u8, value: bson.Value) UpdateError!void { fn set_path(
arena: std.mem.Allocator,
pairs: *std.ArrayListUnmanaged(bson.Pair),
segs: []const []const u8,
value: bson.Value,
) UpdateError!void {
if (segs.len == 1) { if (segs.len == 1) {
if (find_pair(pairs.items, segs[0])) |idx| { if (find_pair(pairs.items, segs[0])) |idx| {
pairs.items[idx].value = value; pairs.items[idx].value = value;
@@ -237,7 +251,11 @@ fn set_path(arena: std.mem.Allocator, pairs: *std.ArrayListUnmanaged(bson.Pair),
} }
} }
fn unset_path(arena: std.mem.Allocator, pairs: *std.ArrayListUnmanaged(bson.Pair), segs: []const []const u8) void { fn unset_path(
arena: std.mem.Allocator,
pairs: *std.ArrayListUnmanaged(bson.Pair),
segs: []const []const u8,
) void {
if (segs.len == 1) { if (segs.len == 1) {
if (find_pair(pairs.items, segs[0])) |idx| { if (find_pair(pairs.items, segs[0])) |idx| {
_ = pairs.orderedRemove(idx); _ = pairs.orderedRemove(idx);

View File

@@ -256,7 +256,13 @@ pub const Reply = struct {
} }
/// Serialize this reply as a full OP_MSG message. /// Serialize this reply as a full OP_MSG message.
pub fn build(self: *Reply, gpa: std.mem.Allocator, request_id: u32, response_to: u32, out: *std.ArrayListUnmanaged(u8)) !void { pub fn build(
self: *Reply,
gpa: std.mem.Allocator,
request_id: u32,
response_to: u32,
out: *std.ArrayListUnmanaged(u8),
) !void {
try write_message(gpa, request_id, response_to, 0, self.pairs.items, out); try write_message(gpa, request_id, response_to, 0, self.pairs.items, out);
} }
}; };