Skip to content

Add a drain-and-restore path for a live serve upgrade (MN-REQ-06.14) - #208

Merged
chouswei merged 2 commits into
masterfrom
cursor/safe-serve-upgrade-416f
Oct 9, 2026
Merged

chouswei merged 2 commits into
masterfrom
cursor/safe-serve-upgrade-416f

Conversation

@chouswei

@chouswei chouswei commented Oct 8, 2026

Copy link
Copy Markdown
Owner

Summary

Patch-level way to upgrade memnet-serve (and the memnet-mcp / product gateway in front of it) without silently dropping a loaded session. Agent usage stays cue, then pin_map, then mutate. Modelled first as MN-REQ-06.14 / MN-VER-06-S12 (SafeServeUpgrade, implemented=true). MN-REQ-06.13 is the concurrent request-isolation cut on master; this branch is rebased onto that.

No version bump and no release.

Design

  • Drain. Admin command memnet admin upgrade-prepare (and envelope {upgrade_prepare:true, admin_token}), gated by MEMNET_ADMIN_TOKEN. Not an agent MCP tool. Quiesce refuses new commands with @ERR: serve_draining|retry_after_s=<seconds>, waits for in-flight work, then snapshots every loaded session through the lossless snapshot path. The manifest (upgrade-manifest.json under MEMNET_STATE_DIR) records session ids, row and edge counts, checksums, serve version, and snapshot format version. An unsaveable session is named (snapshot_unsaveable and any other save failure) and blocks ready_to_stop unless the operator passes --allow-unsaved. A blocked drain does not write the ready manifest.
  • Restore. The new process reads the manifest before it listens, reloads every session, and checks counts and checksums. It prints @STAT: upgrade_restore|ok|n|failed|m and keeps the manifest and snapshot files. Format v1 written by an older patch loads. An unsupported format or a corrupt snapshot fails loudly and does not delete or rewrite the files. Restored sessions keep the same id, ACL bindings, TTL expiry clock, house (session tag map), and product label.
  • Retry. memnet-mcp (TCP) and the product gateway treat serve_draining and a brief connection refusal as retryable, with bounded backoff for MEMNET_UPGRADE_RETRY_S (default 30 seconds). Order: deploy that client tolerance first, then drain and swap the serve. The gateway backend version pin is updated in the same procedure.
  • Helper. memnet-upgrade encodes the side-by-side venv procedure: preflight the new interpreter against the snapshot format, drain, swap the systemd ExecStart from a unit backup, restore, verify, and roll the unit and pin back if verification fails. Operator notes: docs/operations/safe-upgrade.md (rpi5-syson serve 18765 / mcp 18766, the Endleaf engine serve named by the gateway registry, and the droplet gateway).

Downtime that remains

Seconds of connection refusal while the port is closed, covered by the ~30 second retry window. After ready_to_stop, the old process refuses new commands until it is stopped.

Tests

tests/test_safe_upgrade.py and tests/test_sysml_safe_upgrade.py: several sessions (ACL-bound, TTL near expiry, a graph near the row cap); an unsaveable session blocks the drain; a corrupt snapshot fails the restore and leaves the file untouched; client retry across a restart; the same session ids after restore; rollback of the unit and the gateway pin.

Open decisions

  • Neighbourhood reserves are not in the snapshot; re-reserve after restore.
  • A rollback onto the old venv reloads the v1 graph and session id. The old loader rebases TTL from ttl_minutes and ignores the upgrade passport, so ACL must be re-granted if the operator stays on the old venv.
  • Until upgrade-retire, a later restart replays the upgrade snapshots and drops writes made after restore. The helper retires immediately after a clean stat.
  • Composed MCP-in-front-of-gateway retry can stack two 30 second windows. A command timeout is not retried.
  • Quiesce refuses every new non-upgrade command, not only session_open, so the drain cannot livelock.
  • The on-disk manifest names session ids for the operator. Docs and tests do not. The admin usage report still must not.
  • --state-dir passed to prepare must match the unit's MEMNET_STATE_DIR. The helper does not rewrite Environment= lines.
Open in Web Open in Cursor 

cursoragent and others added 2 commits October 8, 2026 17:47
Admin upgrade-prepare snapshots every loaded session, refuses new work
with a retryable serve_draining error, and blocks ready-to-stop when a
session cannot be saved. Startup restores from the manifest and checks
counts and checksums. MCP and the gateway retry across the gap.

Co-authored-by: chouswei <chouswei@users.noreply.github.com>
@chouswei
chouswei marked this pull request as ready for review October 9, 2026 04:00
@chouswei
chouswei merged commit a7013d5 into master Oct 9, 2026
2 checks passed
chouswei added a commit that referenced this pull request Oct 9, 2026
* Bump package identity to 0.19.22.

Hatch, project.toml, changelog, and the version map name the safe serve upgrade path (#208, MN-REQ-06.14). Not 0.20.

* Rewrite the 0.19.22 changelog to the trimmed safe upgrade scope (#210).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants