Skip to content

Peer fleet: two-way serving, heartbeat sync, and live health indicators - #1407

Draft
jeonghun-jj-lee wants to merge 238 commits into
mainfrom
feature/free-tier-fleet
Draft

jeonghun-jj-lee wants to merge 238 commits into
mainfrom
feature/free-tier-fleet

Conversation

@jeonghun-jj-lee

@jeonghun-jj-lee jeonghun-jj-lee commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Transforms the fleet from a hub-and-spoke topology (one server, N thin clients) into a peer fleet where every machine runs its own engine, advertises serving, and maintains live cross-machine roster sync with heartbeat-driven health indicators.

210 files changed, ~31k lines added across 89 commits (all individually PR-reviewed and squash-merged into this branch).


Architecture decisions (ADRs 0023–0029)

ADR Title Status
0023 Base-tier fleet projection Accepted
0024 Fleet thin-client + pluggable transport Withdrawn (superseded by peer model)
0025 Remote-SSH default posture, thin-client as lifeboat Accepted
0026 Generalizable machine capabilities + host-owned roster Accepted
0027 Peer fleet studios + compute federation Accepted
0028 Self-reported device identity Accepted
0029 Unified fleet serving capability Accepted

What shipped

1. Peer fleet topology (ADR 0027 / 0029)

2. Heartbeat + cross-machine roster sync

3. Sidebar fleet section

4. Fleet health / drift popup

5. Thin-client infrastructure (H1 foundation)

6. Fleet setup + enrollment

7. Real-SSH e2e test

8. Session curation (piggyback)

9. Housekeeping


Testing

  • 405 fleet-specific tests (fleet_health, sidebar_fleet_section, sidebar_view) — all green
  • 4166+ total tests passing (the one cli_gate failure is a worktree build-artifact issue, not fleet-related)
  • Real-SSH e2e test: two engines over SSH, p95 = 167ms
  • Manually verified: both machines (Mac Studio + MacBook Pro) show green dots, roster syncs within 30s

Deployment notes

After merge, both machines need a rebuild from latest main to pick up the heartbeat cadence and drift-check changes. Run pnpm sync --fix on each machine.


Issues closed (54)

JJ's fleet issues — implemented (37)

Closes #1258
Closes #1260
Closes #1261
Closes #1262
Closes #1263
Closes #1264
Closes #1265
Closes #1267
Closes #1268
Closes #1269
Closes #1270
Closes #1271
Closes #1272
Closes #1273
Closes #1274
Closes #1275
Closes #1276
Closes #1277
Closes #1278
Closes #1279
Closes #1316
Closes #1318
Closes #1319
Closes #1320
Closes #1321
Closes #1340
Closes #1341
Closes #1342
Closes #1343
Closes #1344
Closes #1345
Closes #1346
Closes #1354
Closes #1358
Closes #1359
Closes #1363
Closes #1375

Aaron's fleet issues — superseded by peer fleet (17)

The peer fleet model (ADR 0029: every machine runs its own engine) makes the hub-client infrastructure these issues target obsolete. The thin-client relay, guard, tunnel, and hub-specific ops are no longer the active architecture.

Closes #780
Closes #792
Closes #651
Closes #775
Closes #776
Closes #781
Closes #782
Closes #1106
Closes #1131
Closes #1227
Closes #1283
Closes #1284
Closes #1288
Closes #1289
Closes #1290
Closes #1294
Closes #1302

NOT closed (future work)

@coderabbitai

coderabbitai Bot commented Sep 21, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@jack-champagne

Copy link
Copy Markdown
Member

holy mother reviewing bits of this now

jeonghun-jj-lee and others added 26 commits September 22, 2026 23:33
…e amicissimo authority (ADR 0023)

The fleet rearchitect (#1106) coupled the never-fork client guarantee to the
amicissimo entitlement: an unentitled machine hits the bootstrap exception
(exit 75), the projection never says role=client, and the guard falls through
to spawning a local server. The base (public) product therefore has no
hardened fleet client.

Add a TS-native base-tier producer: when the amicissimo authority is
unavailable (no entitlement / no checkout) but the machine carries a fleet.json
declaring a real role (client/server), `amico fleet status --projection`
renders + caches a minimal contract-v1 projection itself (mode + topology +
a stable local epoch), instead of exiting 75.

- schema: parseFleetTopology + buildBaseProjection (pure, beside readProjection)
- verb: base-tier at both bootstrap points; enrolled-only (standalone/unenrolled
  keep the exit-75 behavior verbatim — minimal blast radius)
- install.sh: the managed tunnel is CLIENT-only (a server is the tunnel's
  destination, not its client) — no longer dies on a server's missing sshAlias
- one-reader invariant intact: consumers still read only projection.json;
  assert_fleet_guard.sh + the single-parser test stay green

Tests: schema +8, verb +5, installer server-role +1; full fleet sweep 259 green.
Two-slice design for the fleet server-mode Canonical Server on jjs-mac-studio:
Slice 1 launchd reboot-survival service (no auth change), Slice 2 anonymous-on-
loopback + client-only Tailscale tunnel + attach-not-spawn for the Studio editor
(base-tier client role, ADR 0023). Auth decision recorded: loopback + tunnel over
a Tailscale-IP bind with shared token. Refs ADR 0005/0002/0020/0023, PR #1257.
Authored in a parallel Amicode session (fleet thin-client transport track).
Pluggable amicode.fleetTransport provider seam (ssh default / tailscale serve
opt-in / direct); keeps the loopback bind-guard untouched by fronting the
loopback service rather than binding the tailnet IP. Refs ADR 0005/0002/0020/0023,
#792, #775, #777.
… transport ADR

The durable-hub attach/transport/auth design is superseded: #792 (thin-client
relay + inject-auth) and #1260 (fleetTransport seam) own attach/transport/auth;
#955 already ships the anonymous hub runner. Only the reboot-survival residual
survives, now folded into the re-scoped #1258 (held behind #792). Removing this
file also resolves the ADR-0024 number collision — 0024 belongs to the
pluggable-transport ADR.
The client-role never-fork decision was mac-only: isFleetClientGuard
early-returned on process.platform !== "darwin", so a linux/WSL client
silently cold-spawned a local engine (the ADR-0005 split-brain, #1227).

Extract the platform-agnostic decision divertToFleetRelay(binary, state)
into fleet_topology.ts (the role read is already OS-neutral) and drop the
darwin early-return. Never-fork now holds on mac, linux, and WSL alike,
asserted by a non-darwin fixture + a source-guard against the return.
…ck (AC1/AC2/AC6)

A fleet client runs the amicode_service as a relay with NO local engine
(never-fork). Thread a `client` flag through createAmicodeService → the
FleetPlane:

- AC6: suppress the standalone→engine mode flip for a client (getMode stays
  fleet) and emit the relay's OWN named hub-down 503 ("fleet-hub-down" +
  pointer) instead of "engine upstream not available" — there is no local
  engine to be unavailable, and no silent local fall-through. The engine-armed
  machine keeps its existing behavior.
- AC1/AC2 (already-built, now pinned for the client config): UI served locally
  (zero assets cross the WAN), host sessions via the proxied data plane, and
  the credential translated (client mint stripped, hub mint attached).
The client relay refuses to START when its pinned version and the host's
version (GET /global/health → version) disagree beyond a defined tolerance —
an ACTIONABLE message naming both versions and the fix, never a confusing
generic downstream timeout. Protects the host-API-shape assumptions #1262 /
#1264 build on (moved here from #1265 — it is a relay-start gate).

- fleet_version_skew.ts: pure versionSkewVerdict (tolerance exact|minor|major,
  default minor = major.minor must agree), relayVersionGate (start|refuse),
  hostVersionProbe (bounded, honest-null on failure).
- Wired into startAmicodeService: a skewed host returns undefined (no boot);
  matching — or an unreadable host version (deferred to hub-down) — boots.
…n (AC5)

#792 Amendment B's churn objection: an extension-host relay accumulating SSE
upstreams on every window reload (a hub connection storm). Answer it with the
lazy 1:1 pass-through it always was — plus the missing teardown: when the
CLIENT disconnects (a reloaded window closing its stream), destroy the host
upstream. So N reloads re-join each endpoint ONCE, never N× concurrent.

- HubProxy + EngineProxy: res.on("close") → upstream.destroy() (skipping a
  normal writableFinished completion).
- A client-initiated abort sets clientAborted so it never records a
  no-response outcome (a reload must not falsely drive the hub-down posture).

Asserted by a concurrency-counting stub host: N reloads → N joins,
maxConcurrent === 1; the posture stays fleet, no_response_streak 0.
The guard backstop was darwin-only in two places. Make it cross-platform:

- install.sh: enrollment writes the platform-correct VS Code settings path —
  macOS Application Support; linux/WSL ~/.config/Code/User, plus the Remote-WSL/
  SSH server-side ~/.vscode-server/data/Machine/settings.json when present — so
  amicode.opencodeBinary/opencodePort are set on a client of any OS. The guard
  --check host verification runs on every OS now (no non-darwin skip); the
  launchd TUNNEL stays darwin-only (the linux tunnel is #1260). Packaged copy
  kept byte-identical.
- fleet_health.ts: checkFleetGuard/Settings/Role RUN on linux (only
  checkFleetTunnel stays darwin-specific); the aggregate standalone-floor is
  cross-platform.
- extension.ts: the activation-time drift warning fires on linux too (the
  plist read stays darwin-only; the tunnel check self-skips).
…cal dispatch (AC1/AC2/AC5)

In fleet CLIENT mode the host owns all /amicode/* state. dispatch() now
bypasses the ENTIRE local /amicode/* dispatch (the exact-match route table
AND the :262 catch-all) so a REGISTERED route (GET /amicode/problems, POST
/amicode/connections) no longer shadows the fleet branch — the request routes
to the host's authoritative amicode_service (reads via HubProxy, non-GETs via
the write-failure contract). /amicode/fleet/* stays LOCAL (the client's own
posture/mode/staging honesty surface, #1261 AC6). Gated on the client role:
standalone AND the engine-armed base machine are byte-identical.

Extends #1261's stub-host relay test: a registered /amicode/* GET returns the
host sentinel (AC1); the fleet honesty surface stays local (AC1); a mutation
lands + is accepted host-side with the hub mint translated (AC2); standalone
still serves /amicode/* locally (AC5).
…ency pin (AC3/AC4)

The write-failure contract's deriveRefetchPath was session-shaped: an
/amicode/* write fell through to /${segs[0]} = /amicode, dropping the mutated
family and yielding a refetch that hides a partial write. Add
deriveAmicodeRefetchPath — the /amicode/* variant that refetches the FAMILY
GET twin (POST /amicode/connections/credential → GET /amicode/connections),
and dispatch to it from deriveRefetchPath. So a proxied /amicode/* non-GET
resolves through delivered/failed/ambiguous with an /amicode/*-appropriate
refetch, never a raw proxy that hides a partial write (AC3).

Concurrency policy (AC4): LAST-WRITE-WINS with whole-request atomic delivery.
The write pipeline buffers each body fully and sends it in ONE request, so two
clients' bodies never interleave; the single-threaded host applies whole
requests, so the resource holds exactly one client's COMPLETE value. No
lost-update DETECTION (no CAS/ETag) — a client that read-modify-wrote stale
state can silently lose. A two-client concurrent-POST test pins it.
… (AC1/AC2/AC3)

The #1261 relay proxied HTTP + SSE, but both proxies are body-pipe-only and
the relay's HTTP server had no `upgrade` handler — so the integrated terminal
(the engine's WS PTY route, GET /pty/:id/connect) was dead on a thin client:
the handshake fell through to dispatch → the body-pipe HubProxy, which drops
`Connection`, so the host never upgraded (answered 426).

Extend the SAME fleet branch (fleetPlane.client === true && routingMode ===
"fleet") with an upgrade path:

- server.ts: an `upgrade` listener on the relay's http.Server. A fleet CLIENT
  tunnels the upgrade to the host engine via HubProxy.handleUpgrade; every
  other posture (engine-armed base, standalone, no fleet plane) destroys the
  socket — the prior no-listener behavior, so loopback/never-fork is unchanged.
- hub_proxy.ts: HubProxy.handleUpgrade re-implements credential translation for
  the 101/raw-socket path (the body-pipe proxies have no upgrade path to reuse).
  It forwards the engine's OWN 101 verbatim — status line + rawHeaders (the
  engine-computed Sec-WebSocket-Accept) + the head buffer — NEVER a synthesized
  101; keeps Connection/Upgrade/Sec-WebSocket-* so the host upgrade fires;
  strips the client mint + ?auth_token= and attaches the hub mint; preserves
  ?ticket=/?cursor=/Origin. The host's PTY route authorizes by ticket OR Basic.
- Clean teardown (AC3): any end/close/error on either side destroys BOTH sockets
  (pipe with { end: false }) so a dropped WS never leaks the counterpart via a
  half-open FIN; a pending pre-upgrade drop aborts the upstream request.
- server.ts stop(): closeAllConnections() so a live WS/PTY tunnel at shutdown
  can't wedge server.close() on a hijacked socket.

New test fleet_ws_upgrade.test.ts: a WS-echo test against a stub WebSocket
upstream (no `ws` dep — hand-rolled framing). The stub returns a SENTINEL
Sec-WebSocket-Accept a synthesized 101 could never produce, so the client
observing it proves verbatim 101 forward. Covers handshake + both-way frames
(AC1), the PTY route with hub-mint translation + ?ticket=/?cursor=/Origin
preserved (AC2), leak-free teardown on an isolated host (AC3), and the
scope-fence that a non-fleet-client relay never proxies the upgrade.
…undary dedupe (AC1/AC2/AC3/AC4)

The client relay's event stream sent no cursor and the engine's multiplexed
streams emit id: undefined, so a tunnel blip dropped every event in the gap.
The engine's per-session route /api/session/{id}/event?after=<seq> is resumable
(replay durable events after an aggregate seq) but unused.

Make the relay resume it losslessly, client-side (no engine change):
- SessionEventResume tracks the last delivered aggregate seq PER SESSION
  (per aggregate, never one global cursor), carried across relay reconnects.
- HubProxy injects ?after=<seq> on a per-session reconnect and routes the SSE
  body through a byte-verbatim dedupe/track filter (a replayed seq at or below
  the resume point is dropped — idempotent resume, no double-delivery).
- SSE headers are flushed on connect so a fully-caught-up resume still opens
  the stream (Node sends headers lazily; a zero-byte resume would otherwise
  hang the client).
- The non-resumable /event and /global/event streams have no cursor form and
  are LEFT AS-IS (never given ?after=, never assumed resumable) — documented
  out of scope in ADR 0024; making them resumable is separate engine work
  coordinated with #775, not forked from it.

Only armed for a fleet CLIENT (never-fork); the engine-armed base machine
falls back to its local engine on hub-down, so its steady state is unchanged.

Tests (test/fleet_sse_resume.test.ts, reusing #1261's stub-host harness):
AC1 drop-mid-stream loses zero events, gap replayed, boundary deduped;
AC2 per-session cursors resume independently (A@3, B@7);
AC3 /event + /global/event never given ?after=; AC4 steady-state byte-verbatim.
…p — single writer, transition-only (AC1/AC3)
…ed render (AC1/AC2/AC3)

The relay's FleetPostureDetector computes the hub-up-but-slow DEGRADED steady
state the extension's own up/down probe cannot; #780 left feeding it into the
posture state file as an explicit seam. Wire it THROUGH #780's single writer:

- fleet_posture_feed.ts: the RELAY→WRITER seam. Maps the relay posture
  vocabulary (fleet|degraded|standalone) onto #780's PostureFacts and drives
  the SAME FleetPostureStateWriter — no fs write of its own, no second writer.
  Transition-only is the writer's own signature discipline. degraded is
  hub-UP-but-slow (reachable=true), never rendered as fell-back.
- extension.ts: checkFleet's fleet/standalone writes and the manual
  goStandalone fallback both go through the seam (recordPostureState) — one
  fact-builder, the one writer. goStandalone now persists standalone posture
  so the render stops claiming 'attached' the instant the user chooses
  standalone (AC1: never a stale attached claim).
- stack_state.test.ts: locks the honest DEGRADED render (machine, hub,
  reachable-but-slow, timestamped, never as-if-attached) and OS-neutral copy
  (no launchd/plist) — coverage the relay-distinctive state lacked. #780's
  renderPostureLines already handled it; these guard it.

AC1 render for fleet/standalone/missing/corrupt/stale was #780's; this adds
the degraded-state coverage + the relay feed. The client-relay boot that
co-locates the detector with checkFleet is a later #792 world-merge.
The pluggable-transport seam (ADR 0024) as a new module, fully unit-tested:

- FleetTransportProvider interface (the Data Contract the tailscale/direct
  slices consume): resolveBaseUrl() -> URL|undefined, health() -> status.
  undefined base URL = the honest 'no forward bound' (hub-down), never a
  fallback to another provider's URL.
- createSshProvider: the current launchd/systemd -L forward behind the seam.
  resolveBaseUrl() is the loopback hub URL (read late); health() is an active
  probe (reachable+latency+version | unreachable+reason). (AC2)
- resolveFleetTransportKind: the amicode.fleetTransport selector — default ssh,
  independently disableable, a NAMED not-ok for a disabled/unshipped/unknown
  kind, NEVER a silent cross-provider fallback. (AC3)
- transportHealthToOutcome: the SHARED drop-vs-slow-link posture contract this
  slice OWNS — maps any provider's health into the existing FleetPostureDetector
  (reachable -> responded -> p95 decides degraded; unreachable -> no-response ->
  streak decides hub-down); the provider kind rides the reason. (AC4)
- sshForwardArgs + systemdTunnelUnit: the ssh provider's OS forms — the Linux
  systemd form of the launchd plist, loopback-only -L on both ends (ADR 0002/0005).
- hubUrlStringFromProvider: the byte-identical adapter back to the getUrl seam.

Provider-conformance tests reuse the #792 relay stub-hub contract (test/support/
stub_hub.ts). 23 tests.
…no fallback (AC1/AC3)

startAmicodeService constructs the transport provider from the fleetTransport.kind
setting (default ssh) and derives hub.getUrl from provider.resolveBaseUrl() through
the byte-identical adapter:

- unset/ssh reproduces today's launchd-forward behavior exactly — the 13
  existing activation-wiring tests stay green (the no-regression gate). (AC1)
- a not-yet-shipped provider (tailscale/direct) binds NO base URL -> the honest
  'hub upstream not available' 503, NEVER a dial of the configured ssh URL
  (no silent cross-provider fallback). (AC3)
- the boot log names the transport selection outcome (never a silent no-op).

The late-bound hub-URL getter is preserved (a de-armed activation still yields
the honest upstream absence). 3 new end-to-end wiring tests; 16 total in the file.
… (AC1)

- package.json declares amicode.fleetTransport (enum ssh|tailscale|direct,
  default ssh) — the one knob; the description states the no-fallback law and
  that unset/ssh reproduces today's behavior.
- extension.ts reads it (readFleetTransportOption) and passes fleetTransport.kind
  to both startAmicodeService call sites. Empty/unset -> the wiring defaults ssh.
…ifecycle

The Data Contract lists resolveBaseUrl() + health() + start()/stop(). Add the
lifecycle members to the provider interface. The ssh tunnel is OS-managed
(launchd/systemd, installed by the fleet installer) in this build, so its
start()/stop() are honest no-ops that defer to the service manager — the seam
members exist so the tailscale/direct slices (which own their lifecycle)
implement them behind the SAME interface. 24 tests.
…bind preservation (AC tailscale slice)

The host runs `tailscale serve` fronting its loopback service; the client
resolves the host's MagicDNS origin as the base URL. createTailscaleProvider
implements the SAME FleetTransportProvider interface as ssh (resolveBaseUrl →
the MagicDNS origin, read LATE; health() an active probe; start/stop honest
no-ops — serve is host-side, tailscaled is a system daemon). MagicDNS + the
`tailscale serve` mapping are modeled behind injectable seams (no `tailscale`
binary in CI), exactly as the ssh provider models its `-L` forward.

- resolveFleetTransportKind ships `tailscale` in the available set; it is
  independently disableable. `direct` remains the unshipped seam member that
  pins the no-cross-provider-fallback law.
- transportForSelection maps a selection → ITS OWN provider (tailscale → the
  tailscale provider, never a silent ssh substitution); the host wiring uses it.
- tailscaleServeMapping fronts the MagicDNS origin onto a LOOPBACK target
  (127.0.0.1), so the host engine/service bind stays loopback and the mutation
  guard never trips — asserted TOGETHER with MagicDNS resolution in the
  bind-host guard suite, with the rejected direct-tailnet-bind (100.x) contrast.
- health() reuses transportHealthToOutcome (the ONE shared posture contract);
  an unreachable host is the honest hub-down, never a reroute to ssh.
- #775 front-door composition: the serve target is the same loopback front-door
  endpoint the ssh forward reaches, never a bypass.
- Refactor: extracted the shared probeTransportHealth (ssh + tailscale).
amicode-ci and others added 7 commits September 29, 2026 00:51
On the OWNER (Studio) side, archiving a session from a peer landed the
`session.updated` event and updated the per-session store correctly — but
the composer's archived read-only banner only appeared "when you click
around", never off the event alone.

Cause: the composer's `archived` memo read only `get(id)?.time?.archived`.
That key is ABSENT at first render (a live session has no `time.archived`),
so the memo never subscribed to it in the Solid store; when the archive
event ADDED the key, there was no subscription to notify, and the memo
recomputed only when an unrelated dependency (sessionID, a click) changed.

Fix: also touch `time.updated` in the memo — a path that ALWAYS exists and
is bumped on every remember() (the server sets updated=Date.now() alongside
archived in Session.setArchived). That guarantees a live subscription that
fires when the event lands, so the banner flips immediately.

Investigation artifacts kept as tests: two store-content tests in
server-session.test.ts prove the archive event keeps time.archived on the
per-session info store (evict drops only message caches, not info). A
reactivity assertion was attempted but removed — under --conditions=solid
the SSR solid-js build makes store memos inert, so it could neither prove
nor disprove browser behavior (verified with a baseline store probe); the
fix rests on the store idiom, guarded by the content tests + typecheck.
…hive (#1646)

The session-header 3-dot menu (next to compaction) archived the VIEWED
session, then navigated away (splice + evict + navigateAfterSessionRemoval).
The user never saw the read-only banner and had no in-place Unarchive
affordance — the archive appeared to do nothing but move the row.

- archiveSession: stay on the session; drop it from the active list store and
  clear its tab chrome, but do NOT navigate away and do NOT evict (evict drops
  the per-session info the banner + menu toggle read). Force-sync the
  per-session store so the composer banner + menu gate flip at once. Mirrors
  the header + unarchive paths.
- Route both menu archived gates through sessionArchived(id), which touches
  time.updated as the #1646 Solid store subscription anchor so the
  Archive/Unarchive toggle recomputes when the archive event lands.

Tests: message-timeline-archive.test.ts (source-assertion pattern, mirrors
session-header.test.ts #1646). Full timeline suite + server-session #1646
store test green.
Records the message-timeline.tsx content-hash update and the new
message-timeline-archive.test.ts in the overlay provenance manifest
(counts 224→225 app, 907→908 overlay total).
#1646)

Archiving a LOCAL session left the row gone from the active list but the
session still live — no read-only banner, the header menu still showed
"Archive" not "Unarchive". Regression exposed by 9c96c2b, which stopped
evicting/navigating on archive and relies on a force-sync to land the
archived state.

Cause: resolve(id, { force: true }) returned an already in-flight
NON-forced request (`if (pending) return pending`). That request had been
issued BEFORE the archive PATCH committed, so it resolved with STALE
(unarchived) info and re-remembered it, clobbering the fresh archived
state. sessionArchived(id) / the composer memo then read falsy → no flip.
Latent before 9c96c2b because the old archive path evicted + navigated
away, so the stale re-resolve never mattered.

Fix: only coalesce onto a pending request when NOT forced. A forced
resolve always issues its own fetch (and becomes the pending request);
non-forced callers still coalesce as before.

Regression test in server-session.test.ts: force:true re-fetches fresh
archived info even while a non-forced request is pending.
…ate-a-fleet

troubleshoot-fleet (new, public) backs the sidebar Troubleshoot button + boot-time fleet warning that already spawn /troubleshoot-fleet (previously a dangling reference). add-fleet-device (new, public) is the focused single-machine 'add to an existing fleet' path. create-a-fleet is modernized to the two-topology (hub/star + peer studios) implementation: runtime-endpoint truth-probe, one-outcome hub/peer question, end-to-end proof ritual, corrected peer-state statement, and /fleet hand-offs re-pointed to /troubleshoot-fleet. All three carry SSH/Tailscale + cross-OS (mac/linux/WSL) setup guardrails and the honest #1260 Linux-tunnel note.

Skills-only; lint-skills clean (0 errors). No UI wiring or code changes.
…est (#1646)

Follow-up to f54987a. That fixed the resolve()-layer coalescing, but sync() wraps its whole body in runInflight(inflight, id, task), which coalesced unconditionally. The archive read-only flip rides through it as sync(id, {force:true}), so while any non-forced sync/prefetch was in flight for the session the forced task never ran, resolve(force) was never reached, and the composer stayed live — archived row gone from the list but still editable, menu still "Archive". runInflight now takes a force option that bypasses the in-flight coalescing (the inflight twin of the resolve fix); sync() threads it through.

Also migrate the two session-list archive handlers to the in-place model the chat menu uses (#1646): the Rail list (layout.tsx) no longer navigates away, and the Home flyout (home.tsx) force-syncs the per-session store after the list reload — so archiving the viewed session flips it read-only in place instead of bouncing to another session.

Tests: runInflight force-bypass + self-clean (server-session.test.ts); source assertions for the layout + home handlers. Manifest refreshed; drift-gate green.
The message transcript had NO reconnect/resync self-heal: reconcileFromStatus
refreshes only status and the reconnect bootstrap refreshes list/status/global —
none reload the viewed session's messages, and timeline/model.ts re-syncs only
on session-ID change. A silent SSE reconnect or fan-in wedge while staying on
the same session dropped message frames the lastEventID cursor could not replay,
freezing the rail until a manual reload. The entity rail already self-heals
(entity-rail.tsx) via the same signals — the transcript beside it ignored them.

- ui/amicode-entity-view: re-export shouldRefetchOnReconnect/onResync (the
  `@opencode-ai/ui/*` wildcard maps to src/components/*.tsx, so problem.ts's
  predicates need a shim — the app cannot import them otherwise).
- timeline/model.ts: two createEffects mirroring the entity rail — a forced
  session.sync on the disconnect->connect edge and on any resyncCount advance
  (the wedge case, status stays connected). Plain-closure prev, untrack'd
  sessionID guard — track only the signal, no feedback loop.
- session.tsx: thread streamConnected/forceResync from serverSDK().event.
- server-session.ts: harden the orphan-part drop — a genuine orphan (missing
  parent, no active load) now requests a debounced forced resync to backfill
  the parent instead of dropping silently; known-removed/cleared messages still
  short-circuit (no resurrection); per-session in-flight + novel-messageID
  guards collapse a part burst to one sync.

Tests: server-session-orphan-part.test.ts (genuine orphan -> 1 resync; burst ->
<=1; known-removed -> no resync, no resurrection). Timeline effects are glue
over the already-unit-tested predicates. The fourth root of the #1617 staleness
family (after #1636/#1637/#1638).

Non-goals: home.tsx has no transcript to freeze; fleet/remote is parity with
the entity rail's existing limitation, not full per-peer coverage.
@jeonghun-jj-lee

Copy link
Copy Markdown
Contributor Author

Carries the fix for #1646 — the chat message transcript's reconnect/resync self-heal (the fourth root of the #1617 staleness family, after #1636/#1637/#1638).

Commit 005ef265. The transcript now force-syncs on the SSE disconnect→connect edge and on any resyncCount advance (the fan-in-wedge case), mirroring the entity rail; the orphan-part drop is hardened to backfill a lost parent via a debounced resync. Five source files + one test file.

Gates green locally: app-package tsgo clean, engine typecheck PASS (only #1229 base-drift), fast suite 311 files / 5170 tests 0 fail, drift-gate PASS, binary builds (OPENCODE_CHANNEL=dev). Remaining human check: reproduce a live SSE drop in the Ext Dev Host and confirm the rail backfills without a reload.

amicode-ci added 17 commits September 29, 2026 10:34
buildFleetProjection gains an `archived` option; when set, every source is
fetched with ?archived=true and archived peer sessions come back owner-tagged
like active ones. GET /amicode/fleet/sessions?archived=true wires it through.

buildOwnerRoutingProjection unions active+archived so the multiplexer's owner
map can route writes (unarchive/delete) to a peer-owned ARCHIVED session
instead of misrouting to the local engine (#1382 silent-local). Both
OwnerMapFeed closures adopt it.

Part of the peer-fleet feature (PR #1407).
…ns (#1647 S3)

The Archive tab now folds peer-owned archived sessions (from GET
/amicode/fleet/sessions?archived=true, owner-tagged) into the local archived
list via archivedSessionsWithRemote — the composition of the already-tested
peerSessionsFromProjection + mergePeerSessions. Merged on the first page
(bounded recent window); load-more appends the next local page only. A
standalone machine or a fetch failure yields no remote rows, leaving the list
byte-identical to today.

Part of the peer-fleet feature (PR #1407).
…sessions (#1647 S4+S5)

S4: archived rows (home + session-header) now show the owner-machine badge and
control-gate unarchive/delete for remote peer-owned rows — fail-closed to a
lock chip when control isn't held, mirroring the active row. The handlers
guard on remoteDeleteAction so a control-less remote unarchive/delete gives a
clear 'enable control' toast instead of a silent misroute.

S5: the chat 3-dot-menu archive is control-aware (owner/control resolved from
the fleet projection; a control-less remote session gets 'enable control'
guidance rather than a generic failure) and patches with the session's OWN
directory, not sdk().directory — correct for a remote session in a different
project. The #1646 force-sync flip to the read-only banner is unchanged.

Part of the peer-fleet feature (PR #1407).
…ead of asserting instantly

The AC2/AC5 PID-safety tests spawn a real process and assert `ps -o comm=`
sees it as 'opencode' (or that a killed process is dead) immediately. spawn()
returns before the child finishes exec'ing, so the ps-visibility / SIGTERM-reap
windows widen under full-suite CPU contention and the sanity asserts flaked
(green in isolation, intermittent in the 312-file parallel run). Poll with a
timeout instead — same intent (alive AND recognized before boot; dead after
kill), no timing dependency.

Surfaced during #1647; unrelated to that feature.
While the Archived tab is open, refresh it on the same 5s cadence the fleet
projection / owner-map already use, so a local OR remote (fan-in-relayed)
archive/unarchive surfaces without a manual re-open. Remote archive events are
owned by a peer and don't reliably flow through the local directory stores, so
a bounded poll (≤5s) is deliberately chosen over fragile cross-store
reactivity. Cleaned up on tab-close / unmount.

Part of the peer-fleet feature (PR #1407).
…rchive (#1647 S7)

Live-test finding: after archiving a peer-owned session from another machine,
the OWNER (viewing that session) took some time to flip its composer to the
archived read-only banner. Root cause: the #1646 flip is driven by a force-sync
on the ACTING machine and by reconnect/resync self-heal edges; the non-acting
OWNER learns of the archive only via the session.updated SSE event, whose
delivery on that path is jittery — so the flip waited for the next edge/event.

While viewing a session that is NOT yet archived, do a bounded 5s INFO-ONLY
forced resolve (client.session.get → remember; does NOT re-fetch messages, and
is #1646-safe against a stale in-flight). The archived flag converges within
the interval without disrupting an active stream, and the effect self-terminates
the moment archived flips true. Applies to any archive-elsewhere case (peer or
CLI), not just the fleet.

Part of the peer-fleet feature (PR #1407).
…on (#1647 S8)

Live-test finding: archiving a remote session dropped the monitor icon from
its (still-open) tab. The tab's driving indicator — and the composer control
scrim — read the shared control projection, which polled /amicode/fleet/sessions
(active-only). An archived session leaves that projection, so
drivingBannerFromProjection returned null and the icon vanished.

Add ?scope=all to the route: it returns the ACTIVE+ARCHIVED union (reusing S2's
buildOwnerRoutingProjection) with owner AND control overlays on both. The shared
control projection now fetches scope=all, so an open-but-archived remote session
stays resolvable and keeps its icon. Locked with a test asserting the union
carries amicode_control on archived sessions.

Part of the peer-fleet feature (PR #1407).
…1647 S9)

Live-test finding: pressing the archive icon on a remote session in the
Sessions dropdown 'did nothing'. The PATCH did land on the owner, but the
dropdown's active peer rows come from the fleetProjection resource whose
refetch was never destructured or called — so the archived peer row lingered
and the action looked broken.

Destructure refetch from the fleetProjection resource and call it after every
remote mutation (archive / unarchive / delete, active-list and archived-tab),
so the now-archived peer session drops out of the active list immediately (and
reappears on unarchive). Archive also refreshes the Archived tab.

Part of the peer-fleet feature (PR #1407).
…stores the machine tag (#1647 S10)

Live-test finding: when a session was unarchived ON THE OWNER machine, the row
reappeared in this machine's active list but WITHOUT its owner (machine) badge.

The dropdown's active peer rows carry the owner tag from the fleetProjection
resource, which was fetched ONCE on flyout-open and never polled. A remote
unarchive is learned via the SSE event — the row returns from the local store
promptly, but the stale projection had no owner overlay to graft
(mergePeerSessions #1599), so the badge was missing until the flyout was
reopened. The S9 refetch only covers THIS machine's own actions.

Poll the fleet projection every 3s while the flyout is open (matching the
control projection cadence) via a fleetPoll tick folded into the resource key,
so the owner overlay refreshes and the badge returns within the interval.

Part of the peer-fleet feature (PR #1407).
…on create

The per-session multiplexer resolved the owning peer's URL but discarded its
reader token, so the attached-arm create dispatch fell back to the fleet
HubProxy's hub credential. A real peer engine (ServerAuth.authorized) accepts
only its own per-boot password or a token it issued, never the hub token, so
every session create routed to a peer 401'd — and since the machine selector
pre-selects the attached peer, that included creates the user experienced as
local.

Surface peer.token on the reachable ResolvedTarget, add an optional per-request
authOverride to HubProxy.handle (bypassing the hub mint when set), and pass
serverAuthHeader(peer.token) from the reachable create dispatch. No peer token
falls back to the hub mint, preserving the local-engine and single-hub hops
byte-for-byte.

Adds an auth-aware wire test whose peer stub validates the Authorization header
against the peer token — the shipped wire test's stub returned 200 to any
request and could not catch a credential mismatch.
The remote-create fetch wrapper rebuilt the request headers from only
init.headers when it armed x-amicode-owner, which shadowed the client's base
Authorization header (createOpencodeClient applies the Basic credential as the
client's base headers, not per-request init). The armed path-less POST /session
therefore went out with no credential and 401'd at the service gate, while every
other call kept the base credential.

Extract composeArmedRequestHeaders, which seeds the base Authorization before
arming the owner header and rewrites the request only when armed; unarmed calls
are left untouched so the client base headers still apply. Adds unit coverage
for the armed-create, unarmed, and non-create cases.
On a base-tier machine (role=server, observationOnly — no premium multiplex
plane) a path-less POST /session picking a peer fell through to the local
engine, so "create a session on the Studio" ran locally. The observation path
delivered observe + control over existing sessions but never remote creation.

Add ObservationWriteRouter.resolveCreate: the create mirror of resolve, keyed
off the x-amicode-owner header (a create has no session id yet), authorized by
the same control-grant write gate and routed to the owner with the peer reader
token. Dispatch consults it beside the pathed-write resolve.

Close the read-after-create gap ("This session cannot be found"): the new
session is not yet in the pull-only projection when the app reads /session/{id}.
SessionOwnerMap gains pending create-time bindings (recordPendingOwner) that
survive projection refreshes, are superseded once the projection includes the
session, and expire after a TTL. proxyCreateToPeer proxies the create buffered,
reads the new session id from the response, and records the binding so the
immediate follow-up read routes to the peer.

Tests: resolveCreate unit + dispatch integration (create -> peer with reader
token; no-owner -> local; no-grant -> 403), pending-binding unit (immediate /
survives-refresh / superseded / TTL), and a full create->ownership-recorded->
read-routes-to-peer integration. 182 across the touched service suites; tsc clean.
A free-tier remote session create is proxied to the owning peer with the
creating machine's ?directory=<creator project> forwarded verbatim. The peer
engine then files the session under that foreign path, but the peer's own
sidebar lists session.list({ directory: <the peer's open project> }) — so the
remote-created session was invisible on the machine that actually runs it.

Drop the directory query in proxyCreateToPeer so the peer resolves its own
default project scope. The follow-up read routes by session id (owner map),
independent of directory, so nothing downstream needs the creator's path.
…on the peer

Slice 1 stripped the directory unconditionally, which fixed the foreign-path
case but made a proxied create fall back to the peer's throwaway ambient temp
cwd — a project the peer's UI never lists, so the session stayed invisible.

Match-by-path instead: probe the owning peer's GET /project and KEEP the
creator's directory iff it is a worktree there (the common shared-repo case,
so the session lands in the repo the peer's sidebar shows); otherwise strip it
to the peer's default scope. Probe failures fail safe to strip.
…fleet projection

The fleet Sessions surface (GET /amicode/fleet/sessions) is cross-project by
design, but its LOCAL source bailed with present:false 'engine mint not armed'
whenever no per-boot Basic password was threaded to it — even though the local
engine is an open loopback (the EngineProxy already dials it credential-less).

Effect: a machine could not see its OWN sessions in its fleet dropdown. The
creating machine got away with it (it reads the owning peer with a token); the
OWNING peer, whose sessions are local, showed nothing.

Fetch the local source WITHOUT an Authorization header when no password is
configured (allowUnauthenticated on the local call only). Peers/hub keep the
old no-credential bail. A real 401 is still named 'unauthorized' honestly.
The Sessions dropdown built its LOCAL rows by iterating only the known project
directories (projects.list()) and folded in ONLY remote rows from the fleet
projection (peerSessionsFromProjection drops every is_local entry, assuming the
local list already has them). A local session in a directory the store never
iterates — a remote-created session that landed in the owner's ambient temp cwd,
or any unlisted project — fell through BOTH paths and was invisible on the
owning machine's dropdown, even though the cross-project projection contained it.

Add allSessionsFromProjection (local + remote) and feed the active dropdown from
it; mergePeerSessions still dedupes with the local store winning, so listed-dir
rows are unchanged and only the missing local sessions are surfaced. The dropdown
now shows all sessions the projection knows, cross-project and cross-machine —
independent of which directory a remote create happened to land in.
The rail's running dot froze while a foreground subagent (Task tool) ran,
settling to done until a reconnect/resync/switch. The dot is driven by
turnIsRunning = isActive && status === "busy" (rows.ts) where status is the
RAW session_status[parent] leaf — it never consulted the floored session_working
that workingTurn reads. Both existing floors are inapplicable to a foreground
subagent: turnActive never arms (the engine has no session.execution.started
publisher) and streamActiveParts arms only on the parent's own part.delta, but
the parent turn blocks in the task tool on background.wait(child) emitting no
parts while the child streams on its own session id. So the parent's liveness
hung on the raw busy leaf, which a stray/reconcile idle (temporal-guard bypass
during a long subagent) blanked.

Fix — an ancestor-of-active-child floor, pushed as a REACTIVE store leaf (the
diff_version propagation pattern), NOT a plain-Map read (which would not
re-trigger the parent's projection memo on a first child spawn — the trap the
deliberation caught):

- server-session.ts: new store leaf session_child_active (count of non-idle
  descendants per ancestor). refreshChildActive(child) recomputes the child's
  0/1 contribution against a remembered value and applies the DELTA up the
  taskSpawnParent chain via setData (idempotent — a duplicated/missed frame
  cannot drift the count). Called on recordTaskSpawn (handles child-before-
  mapping ordering), session.status, the execution terminals, and session.error.
  session_working ORs (session_child_active[id] ?? 0) > 0.
- evict: lowers the session's contribution to ancestors BEFORE severing links,
  then drops taskSpawnParent (self + children), childActiveContribution, and the
  session_child_active leaf — closing the pre-existing taskSpawnParent evict leak.
  session.deleted now delegates that teardown to evict (its manual deletes would
  have severed links before the floor could be lowered).
- projection.ts / rows.ts / message-timeline.tsx: thread a floored
  accessor (session_working) into the projection; turnIsRunning uses it. status
  stays the RAW leaf for the retry row and the tailStreaming check.

Tests: server-session-child-active.test.ts (6) — parent held across a stray
reconcile idle while a child is busy; floor lowers on child idle / exec terminal;
idempotent under duplicate busy frames; the reactive SOURCE (the store leaf)
transitions on child edges; evict clears the index. (bun resolves solid-js to
its server build so createMemo does not track — the end-to-end re-render is the
human live-reproduce; the store-leaf write is the reactivity mechanism.)

The fifth root of the #1617 staleness family (after #1636/#1637/#1638 and the
#1646 transcript self-heal). Client-side only; no engine change.
@jeonghun-jj-lee

Copy link
Copy Markdown
Contributor Author

Carries the fix for #1649 — the thought rail's foreground-subagent liveness (the fifth root of the #1617 staleness family).

Commit 4b4dbf21. The rail's running dot was driven by the RAW session_status[parent] busy leaf, which a foreground subagent leaves un-refreshed (parent blocks in the task tool; neither the turnActive nor streamActiveParts floor applies) and a reconcile can stray-idle. Fix adds a reactive ancestor-of-active-child floor (store leaf session_child_active, propagated like diff_version) and points the rail's turnIsRunning at the floored session_working instead of the raw leaf. Client-side only.

Gates green locally: app tsgo clean for the 4 touched files (one pre-existing unrelated error on this branch at session-composer-region-controller.ts:144); fast suite 314 files / 5198 tests 0 fail; drift-gate PASS (913 files); binary builds (dev/1.18.30). New tests server-session-child-active.test.ts 6/6. Remaining human check: launch a subagent and confirm the rail stays live through the child run.

amicode-ci added 3 commits September 29, 2026 22:14
… remote delete

The titlebar Sessions flyout had three defects on remote (peer) rows:

- Laggy list: the <For> rendered every loaded session across all projects
  plus all fleet peers; `activeLimit` only bounded the backend fetch, never
  the DOM, so hundreds of rows each carried the per-row
  useSessionTabAvatarState subscriptions and re-diffed on every 3s fleet poll.
  Slice the rendered list to `activeLimit()` (after the search filter, so
  search still spans everything loaded) via `pagedActiveSessions`, and
  redefine `hasMoreActive` to mean "rows exist beyond the slice, or the fetch
  cap was hit". Caps the DOM at 50 rows (+50 per Show more).

- Archive overlaid the machine tag: the owner tag now fades out on
  hover/focus so the right-edge affordance never covers it (active + archived
  remote rows).

- Delete on active remote rows: removed the owner-routed remote-delete
  arm→confirm + its plumbing (onRemoteDelete, remoteDeleteSession). Remote now
  matches local — archive from the active row, delete from the Archived tab
  (deleteArchivedSession is already owner-routed + control-gated).

Also fix a latent bug the overlay typecheck gate surfaced:
session-composer-region-controller.ts called `resolve` on the
directory-scoped session store, which only re-exported `sync`. At runtime the
#1647 S7 owner-side archived-flag convergence threw every 5s (caught/ignored)
and never fired. Add an info-only `resolve` passthrough to directory-sync.ts
delegating to serverSync.session.resolve (NOT `sync`, which force-reloads
messages and would disrupt a live stream).

Manifest refreshed for the three overlay files (drift gate verifies).
Two defects on tagged (peer/owner) rows in the titlebar Sessions flyout:

- Hover flash: while the flyout is open, the 3s fleet poll re-fetches the
  projection and mergePeerSessions mints a NEW object for every peer/owner-
  tagged row on each fetch. SolidJS <For> keys by referential identity, so a
  new reference at an index disposes + remounts the row — resetting its CSS
  :hover state and replaying the machine-tag opacity transition under a resting
  pointer. The archived tab hits the same on its own 5s poll. Add a pure
  stabilizeSessionIdentity(next, cache) helper: an id-keyed cache that returns
  the PREVIOUS row object when a row's render signature is unchanged, so
  unchanged rows keep their reference (and DOM node) across the poll. The
  signature (sessionRowSignature) excludes time.updated — it churns per streamed
  token but changes nothing the row paints — and tracks title/directory/owner/
  control. Wired into the active (activeIdentityCache) and archived
  (archivedIdentityCache) memos feeding <For>; component-scoped so it survives
  memo re-runs; self-evicts departed ids.

- Owner tag vanished on hover: instead of fading the machine tag to nothing to
  reveal the action cluster, teleport it — the resting badge still fades out,
  and the SAME tag reappears inside the (fade-in) action cluster to the left of
  the archive/delete controls. Reads as the tag sliding aside for the action,
  not disappearing. Applied to the active row (local + remote clusters) and the
  archived row; marked data-slot="…-teleported".

Tests: 5 stabilizeSessionIdentity/signature unit tests (same-ref-on-unchanged,
new-ref-on-content-change, time.updated is ignored, order + eviction) and
source assertions for the teleported badges + the identity wiring. Manifest
refreshed for the four overlay files (drift gate verifies).
…n hover

Follow-up to 62e45b6, which fixed the wrong way: it faded the resting tag out
and rendered a SECOND, max-w-[40%]-truncated copy inside the archive's own
absolute cluster — so the tag AND the archive appeared to jump into place, and
the copy was shortened.

Now there is ONE tag element that stays full length and slides left by exactly
the action cluster's rendered width, with the archive pinned at the row's right
edge (it never moves). The cluster is no longer an absolute right-edge overlay;
it reveals IN FLOW via a CSS-grid `grid-cols-[0fr]→[1fr]` column, so it reserves
the controls' exact width — the row button (now flex-1, was w-full) shrinks and
pushes the tag over by precisely that amount. Width-agnostic: no hardcoded pixel
shift (the large icon-button width lives in a compiled stylesheet). The tag is
never faded and never truncated on hover; the title gives up the space instead.
Applied to the active row (local + remote clusters) and the archived row.

The #1652 identity-cache flash fix from 62e45b6 is unchanged.

Tests updated: one owner badge per row (no teleported copy), the tag is not
faded on hover, and the cluster reveals via the 0fr→1fr grid (the old absolute
overlay is gone). session-header 18/18; fleet-peers/dropdown/archived 100/100;
overlay typecheck + drift green. Manifest refreshed.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment