Repository navigation
Add a drain-and-restore path for a live serve upgrade (MN-REQ-06.14) - #208
Merged
Merged
Conversation
Admin upgrade-prepare snapshots every loaded session, refuses new work with a retryable serve_draining error, and blocks ready-to-stop when a session cannot be saved. Startup restores from the manifest and checks counts and checksums. MCP and the gateway retry across the gap. Co-authored-by: chouswei <chouswei@users.noreply.github.com>
chouswei
marked this pull request as ready for review
October 9, 2026 04:00
This was referenced Oct 9, 2026
chouswei
added a commit
that referenced
this pull request
Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Patch-level way to upgrade
memnet-serve(and the memnet-mcp / product gateway in front of it) without silently dropping a loaded session. Agent usage stays cue, thenpin_map, thenmutate. Modelled first as MN-REQ-06.14 / MN-VER-06-S12 (SafeServeUpgrade,implemented=true). MN-REQ-06.13 is the concurrent request-isolation cut on master; this branch is rebased onto that.No version bump and no release.
Design
memnet admin upgrade-prepare(and envelope{upgrade_prepare:true, admin_token}), gated byMEMNET_ADMIN_TOKEN. Not an agent MCP tool. Quiesce refuses new commands with@ERR: serve_draining|retry_after_s=<seconds>, waits for in-flight work, then snapshots every loaded session through the lossless snapshot path. The manifest (upgrade-manifest.jsonunderMEMNET_STATE_DIR) records session ids, row and edge counts, checksums, serve version, and snapshot format version. An unsaveable session is named (snapshot_unsaveableand any other save failure) and blocksready_to_stopunless the operator passes--allow-unsaved. A blocked drain does not write the ready manifest.@STAT: upgrade_restore|ok|n|failed|mand keeps the manifest and snapshot files. Format v1 written by an older patch loads. An unsupported format or a corrupt snapshot fails loudly and does not delete or rewrite the files. Restored sessions keep the same id, ACL bindings, TTL expiry clock, house (session tag map), and product label.serve_drainingand a brief connection refusal as retryable, with bounded backoff forMEMNET_UPGRADE_RETRY_S(default 30 seconds). Order: deploy that client tolerance first, then drain and swap the serve. The gateway backend version pin is updated in the same procedure.memnet-upgradeencodes the side-by-side venv procedure: preflight the new interpreter against the snapshot format, drain, swap the systemdExecStartfrom a unit backup, restore, verify, and roll the unit and pin back if verification fails. Operator notes:docs/operations/safe-upgrade.md(rpi5-syson serve 18765 / mcp 18766, the Endleaf engine serve named by the gateway registry, and the droplet gateway).Downtime that remains
Seconds of connection refusal while the port is closed, covered by the ~30 second retry window. After
ready_to_stop, the old process refuses new commands until it is stopped.Tests
tests/test_safe_upgrade.pyandtests/test_sysml_safe_upgrade.py: several sessions (ACL-bound, TTL near expiry, a graph near the row cap); an unsaveable session blocks the drain; a corrupt snapshot fails the restore and leaves the file untouched; client retry across a restart; the same session ids after restore; rollback of the unit and the gateway pin.Open decisions
ttl_minutesand ignores the upgrade passport, so ACL must be re-granted if the operator stays on the old venv.upgrade-retire, a later restart replays the upgrade snapshots and drops writes made after restore. The helper retires immediately after a clean stat.session_open, so the drain cannot livelock.--state-dirpassed to prepare must match the unit'sMEMNET_STATE_DIR. The helper does not rewriteEnvironment=lines.