Skip to content

refactor: replay the AOF through the keyspace - #416

Merged
kacy merged 1 commit into
mainfrom
refactor/replay-through-keyspace
Sep 24, 2026
Merged

kacy merged 1 commit into
mainfrom
refactor/replay-through-keyspace

Conversation

@kacy

@kacy kacy commented Sep 24, 2026

Copy link
Copy Markdown
Owner

ember-persistence/src/recovery.rs replayed the AOF against its own HashMap<String, (RecoveredValue, ttl)>, with a hand-written copy of every write command's effect: INCR, LPUSH, LTRIM, ZADD, HINCRBY, vectors and the rest. That was about 560 lines of command logic outside the keyspace, and it had already drifted. For example, an EXPIRE record could bring back a snapshot key whose TTL had run out, because the map kept expired keys until the end.

Now:

  • ember-persistence only reads files. recover_shard hands snapshot entries to Recover::restore and AOF records to Recover::apply.
  • The shard implements Recover with its keyspace (shard/persistence.rs). A snapshot entry goes through Keyspace::restore. An AOF record becomes the request that wrote it, via from_aof_record from refactor: move the AOF record to request converter into ember-core #415, and runs through dispatch, the same path as a live command.
  • Because nothing can be undone once it's in the keyspace, the AOF is read twice. The first pass decodes every record without applying any, trims a torn tail, finds the checkpoint where replay starts (fix: skip an AOF that is older than the snapshot #408), and detects a stale AOF. A file that is corrupt partway through is skipped entirely, as before. The second pass applies the records.
  • Schema registrations are replayed even from before the start point or from a stale AOF, since schemas live only in the AOF.
  • Memory limits are off during replay and restored afterward. Recovered keys were all stored once already, so none should be refused or evicted.
  • RecoveredValue, RecoveredEntry and the shard's own RecoveredValue to Value conversion are removed.

Tests:

  • The persistence tests now check what the driver hands over: snapshot and AOF order, expired snapshot keys, a corrupt snapshot, mid-file corruption, a torn tail, checkpoints, and schemas from a stale AOF.
  • Command behavior during replay is tested in ember-core against a real keyspace, including the memory limit.

About 1,200 fewer lines in total.

Recovery repeated the effect of every write command against a HashMap
in ember-persistence, a second copy of the command logic that could
drift from the keyspace. ember-persistence now only reads the files and
hands entries and records to a Recover target. The shard implements it
with its keyspace: each record becomes the request that wrote it and
runs through dispatch. The AOF is scanned once before anything is
applied, so a file corrupt mid-way is still skipped as a whole.
@kacy
kacy merged commit 3f2e6f6 into main Sep 24, 2026
12 checks passed
@kacy
kacy deleted the refactor/replay-through-keyspace branch September 24, 2026 21:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant