This audit describes the initial format, not power-loss qualification. Package versions, source commits and license identifiers are recorded in dependency-survey.json. The server uses RF3 from its first implementation; the embedded adapter is used for local testing, administrative tools and benchmarks.
ZoneTree 1.9.8 is the ordered materialization. Its data WAL is synchronous and compression is disabled in the initial profile. KeyLoad's checksummed commands.wal is authoritative for recovery: it stores the exact compiled mutations, including the outcome, apply watermark, indexes and queue state. Recovery does not reevaluate preconditions against a partly materialized transaction.
The adapter compiles each operation under a write gate, writes a bounded frame, calls Flush(true), applies all mutations to ZoneTree and releases the gate. Readers hold a read gate across their point/index scans and projection. A failure after append makes the open adapter unusable until recovery. An incomplete final frame is truncated; a complete frame with a checksum or sequence error fails closed. Only a successful recovery permits service to resume.
Staging accounts canonical base64 frame bytes incrementally, including replacements and deletions, before copying additional mutations. Compilation also validates the exact serialized frame before any journal write and reuses that prepared payload during commit. If domain/index/outbox expansion exceeds the frame limit, the state machine resets staged effects and persists a bounded ResourceExhausted outcome, clock and apply position instead. The rejection therefore advances Raft apply and survives retry/reopening. Nullable array elements and undefined mutation/catalog enums are rejected before dereference. A mutation's declared resource must match its canonical collection/stream/queue/graph owner before authorization; shadowed JSON properties cannot select a different grant. Malformed client data cannot become an unapplicable replicated entry through those validated paths.
The separate node ownership lock prevents concurrent writers. Returned bytes and staged mutations are copied. Identity and backup files are versioned; the backup manifest verifies every canonical file with SHA-256. A backup is made under the apply gate and reports that exact cut. Restore changes the incarnation, invalidates old signed tokens, resets the Raft/Orleans watermark and membership, and explicitly persists paused dispatch even if the old journal had unpaused it.
Format 2 adds a checkpoint prefix with ordered live records, per-frame SHA-256 checksums and a complete-image checksum/count. It records both the local journal position and the Raft apply position. Installation verifies the entire incoming image before closing the old generation, flushes the staged file and atomically replaces the canonical journal. The derived ZoneTree generation is rebuilt. Installation keeps the node identity, increments the read generation and invalidates existing query cursors; offline compaction preserves the logical cut and cursor generation. Process-kill regressions cover snapshot write/flush and the two publication boundaries, including reclamation after reopening.
The .NEXT 6.8.1 persistent Raft log owns terms, votes, log matching and consensus commit. DurableRaftLog requires native manual checkpoints (FlushInterval = Timeout.InfiniteTimeSpan), serializes its public append/commit paths and awaits the native foreground FlushAsync before acknowledgement. Background flush notifications do not establish this boundary for an isolated first append or snapshot installation. Four regressions invoke append through the actual IPersistentState interface, kill the process immediately after ACK and reopen the tail using private page memory; they do not rely on graceful disposal or shared memory pages surviving exit.
The state machine produces native Raft snapshots from a verified canonical store cut. A snapshot-only regression verifies that installation acknowledges without a subsequent append and reopens at that cut. The RF3 test erases one stopped follower's data directory, requires native snapshot installation, checks the fresh node identity and read generation and verifies documents, retained events, subscription checkpoints, inbox receipts and command outcomes through the HTTP API. Incoming bodies are authenticated and rewound for both Body and BodyReader; outgoing one-shot payloads are serialized once into a bounded disk spool. Raft term, snapshot metadata, request ID and content type are included in the peer signature.
Incoming installation additionally flushes and atomically publishes a bounded, scoped intent before entering the native provider. The provider can publish a file on a failed write; the guarded state-machine lifecycle removes an unverified attempt while preserving an image whose canonical installation has begun. Startup resolves the intent before restoring native snapshots or opening the WAL and reclaims unpublished temporary files. A corrupt or missing image whose cut is already materialized, malformed intent or established snapshot corruption fails closed and preserves evidence. Four transfer regressions cover interruption, cancellation, checksums and foreign scope; five actual process kills cover intent, body writing, rejected-file publication, verified-image and completed installation boundaries. See docs/design/replica-snapshots.md.
The bounded leader writer replicates a trusted operation with one leader-chosen evaluation time. It returns success after consensus commit and local state-machine apply, then resolves the persisted outcome under current authorization. Cancellation or response loss after admission has an unknown write outcome; retries must retain the command ID. A minority cannot acknowledge writes.
Per-source subscriptions journal bounded delivery references, contiguous checkpoints and signed lease generations. Processing effects, inbox receipts, ACK and the command outcome share one compiled redo frame. Seven process-kill scenarios reopen the real store before retrying a processing command and require the effect and checkpoint to agree; a second inbox retry must not apply the effect twice. Seek is an explicit generation fence. Retained topic quotas fail the producer batch before publication; reclamation and cross-partition coverage are separate pending work.
The committed system outbox stores accepted mutations and document before/after images alongside their effects. A generation-bound projection batch commits canonical effects, its replay receipt, checkpoint and resulting outbox entries atomically. A separate bounded progress reserve permits one reserve-using batch per consumer until the retained prefix advances. The reservation fence is in that same atomic frame; ordinary producers cannot spend it. Seven additional process-kill scenarios use a full producer outbox and require checkpoint, reservation fence, effect, outbox tail and command outcome to agree after reopening, then replay without duplicating the effect. Active projection and rebuild consumers pin retention. Public document feeds and scalar live-query deltas use current row/field permissions; ACL changes fence old cursors. The initial live snapshot and captured tail share one read gate.
Strong reads on a leader force a quorum barrier tied to that leadership term. Followers request this barrier through an authenticated peer endpoint and wait for local application of the returned committed position. The provider's follower synchronization API alone is not used as KeyLoad's strong-read acknowledgement. Text and vector search, SQL predicates, row authorization and returned field projection each stay within one storage read gate.
Search and graph reads charge raw scan keys/values and all document/edge point dereferences against one cumulative read budget. Hybrid modalities share that budget. Cancellation and a monotonic deadline bound execution between operations. Graph visit caps include filtered/hidden adjacency work across vertices; serialized result caps include metadata. Exact BM25 retains query-term counts with a corpus-wide token ceiling rather than retaining every corpus word. These read bounds do not change canonical writes or establish a total native/RSS limit; see docs/design/bounded-reads.md.
The cluster coordinator reserves bounded command count and retained payload bytes before enqueueing, scoped by verified tenant/principal. Short delivery, membership and dispatch commands have a separate control reserve and bounded priority. Accepted commands retain their reservation across caller cancellation until the worker finishes; shutdown releases queued reservations with an unknown-outcome response. Local admission configuration does not alter canonical apply decisions. An RF3 test exhausts data admission and verifies unpublished catalog writes, unclaimed rejected IDs, continued dispatch commits and ready Orleans routing on all voters.
Public HTTP requests reserve node capacity before deserialization and quorum reads; authenticated requests additionally reserve current verified tenant/principal counts. Query/search/graph and bulk-read routes share modeled working-space admission. A separate public control pool permits delivery and dispatch while data admission is full. Kestrel bounds declared and chunked request bodies; oversize protocol requests return typed ResourceExhausted without claiming the command ID. Reservations last through response processing. Unit and RF3 tests qualify those admission paths; total native allocations, process RSS and disk budgets remain pending qualification.
Orleans membership is a replicated catalog record, bootstrapped through consensus before Orleans starts. Transient quorum/admission failures during silo startup are retried with fresh host state while the native Raft endpoint remains available, so a voter restart during election does not terminate its process. Shutdown cleans up partially started hosts. Grains route commands; they do not own files or durability. The first topology contains one physical shard and many separately scoped atomic partitions.
Local tests run on macOS arm64 with .NET SDK 10.0.401. The recovery suite executes 1000 seeded real-process kills before/after journal flush and at partial materialization points, verifies atomic recovery, rejects complete-frame corruption and restores a verified backup. The RF3 Aspire suite owns three independent node processes, kills the leader, retries a committed command, checks inbox/effects/ACK, rejects minority operations and restarts voters.
The local result and crash-stage distribution are recorded in kernel-qualification.json. Each trial writes its seed, fault stage, mutation index, recovered values and platform to artifacts/qualification/crash-trials-*.jsonl; CI retains these as run artifacts.
The advertised profiles remain ProcessDurable and QuorumProcessDurable. These tests do not establish filesystem directory-entry persistence, storage-controller guarantees, real power-cut behavior or every operating system/filesystem combination. Linux/macOS/Windows CI and broader network faults are tracked separately. Large snapshots under transport deadlines, automatic canonical compaction, power-loss qualification, network partitions, long histories, shard movement and the 72-hour endurance gate remain required for the broader architecture's durable release profile.