Skip to content

Idle-suspend blocked tasks on remote hosts, and resume them there - #828

Merged
bborn merged 5 commits into
mainfrom
task/5717-idle-suspend-blocked-tasks-on-remote-hos
Oct 8, 2026
Merged

bborn merged 5 commits into
mainfrom
task/5717-idle-suspend-blocked-tasks-on-remote-hos

Conversation

@bborn

@bborn bborn commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

Why

On 2026-10-08 ik-agents ran out of memory (680 MB free of 15 GB, swap full, load 245): 40 blocked tasks placed there by a coordinator still had their claude agent plus puma/webpack/sidekiq running, some for 20 days, although idle_suspend_timeout was 24h.

suspendIdleBlockedTasks finds an agent through the local tmux and ps (getClaudePID). A placed task's window is in task-daemon-remote-<coordinator> on the host, so it reads as PID 0 and is skipped. ty sessions cleanup is local-only too. #823's per-host sweep only handled done/archived tasks.

⚠️ On deploy

The first sweep after this ships will suspend every overdue blocked placed task at once: every blocked task on every host that has been idle past idle_suspend_timeout (and whose panes show no recent activity). On ik-agents that is the ~40 tasks from the incident, all in that one sweep. That is the intent, but expect a burst of "Agent suspended on " log lines, and of SIGTERMs to their dev servers.

What

Remote idle suspend (remote_idle_suspend.go, wired into #823's sweep, now sweepRemoteSessions). One ssh call per host per sweep carries both the finished and the idle work, with the same backoff for unreachable hosts. The sweep runs from the daemon loop, so a status change never waits on it. Nothing is deleted.

Windows. For a blocked placed task whose blockedIdleDuration ≥ idle_suspend_timeout, the host script lists this coordinator's session with #{window_id} #{window_name} #{window_activity} #{pane_start_command} and ends task-<id> / task-<id>-shell by exact name and window ID. It does not end them, and does not touch the task's side processes either, when:

  • a window saw activity within the timeout. Someone is typing in it, whatever the database says. The task isn't asked about again until it could next be idle, so a busy pane costs no ssh call per sweep.
  • the agent window belongs to another run. That is the case when its start command names a WORKTREE_RUN_ID other than the run being suspended, e.g. a reply resumed the task between the sweep's read and its ssh. I match on the run ID string rather than on WORKTREE_RUN_ID='<run>' literally, because tmux re-quotes #{pane_start_command}. Windows from before run IDs existed count as the run's own. A window naming another run is spared only while it could be a newer run's: afterwards the sweep re-reads the stored run, and if it still equals the sweep's run, the window is an older run's and the next sweep (30s later) ends it and suspends the task.

Window names are always sent. daemon_session only decides whether the task needs asking.

Side processes. These are ended in the same call once reap_blocked_idle has also passed; with that setting disabled they are left alone. The host is shared (on ik-agents ty is the same agents user as the host's own TaskYou), so this part is deliberately narrow:

  • It only looks at this user's processes (ps -u "$(id -u)" -ww -o pid=,ppid=,args=).
  • It never signals a coding agent (claude, codex, …) or tmux.
  • It never signals anything with a live tmux pane as itself or an ancestor. Panes come from list-panes -a on every tmux socket the user has, plus the default one, walking up the ppid chain as reaper's hasLiveAncestor does.
  • It never signals the sweep's own shell, its ancestors, or its children.
  • It matches only the recorded worktree. On Linux a process matches when /proc/<pid>/cwd is inside it, which also catches sidekiq, foreman and bin/dev. Elsewhere (no /proc) it matches by the full worktree path, or a puma title [<dir>], with a character boundary after either.
  • There is no <id>- fallback. A task with no usable recorded worktree has its windows ended and its side processes left running, and its log says so.
  • Each process gets SIGTERM, then SIGKILL 2s later.
  • A host without awk or ps is an error (retried with backoff), not "nothing to end".

State. It records what a local suspend records: the tmux placement (daemon_session, window and pane IDs) is cleared, the task stays blocked, and the log says Agent suspended on <host> (idle timeout); session preserved for resume. ty sessions list and the attach hint stop showing it as running. The remote run row is forgotten once nothing more will be done for it.

Resume (remote_session.go). Replying to a suspended task goes through retry. A placed task's retry used to start a fresh Claude with the reply appended to the prompt. Placement is sticky and the remote worktree is reused, so now:

  • A placed Claude run starts with --session-id <uuid>, and ty records that ID.
  • A retry looks the session up on the host. It uses the recorded ID if that transcript is still there, else the newest *.jsonl in the worktree's Claude projects dir (for runs launched before this change), and launches claude --resume <id> with the reply.
  • If ty cannot find out, because the host is unreachable, the lookup times out or the worktree is missing, the run fails with "Could not reach to resume this task's session" and the task is blocked. It never silently starts a new conversation.
  • If the host has no session at all, the task starts fresh as before, and the log says so.
  • ty retry --replace (ClearTaskPlacement) forgets the recorded session ID when the old placement was a remote host. Otherwise HasLocalState would read it as a local first attempt and pin the task to this Mac without asking the resolver. A local placement's session is kept.

Settings: reuses idle_suspend_timeout and reap_blocked_idle; nothing new.

Tested

  • TestIdleBlockedRemoteTasksAreSuspended uses fix: end a placed task's remote agent session when it is done or archived #823's fake ssh plus a file-backed fake tmux (now with window activity, start commands and list-panes), and a fake ps that lists only real tail -f processes the test starts.
    • Ended: the idle task's windows (including a pre-run-ID agent window), its dev server and its puma. On Linux also a sidekiq that only has the worktree as its cwd.
    • Left alone: another task's claude whose prompt names the worktree; a tmux server whose command names it; a process under a live pane; another coordinator's task with the same ID; a task ID that only shares a prefix (13 vs 1); [1-Oct-2026] text; every side process of a task with no recorded worktree; recently blocked and processing tasks; a task whose pane was typed in a minute ago; a task whose agent window is another run's; same-ID windows in other sessions; on macOS, the sidekiq.
    • Also checked: no files change; the log lines (including "side processes … left running"); a second sweep makes no ssh calls; the unreachable host backs off and is retried.
    • I removed each guard in turn (agent/tmux exclusion, pane ancestry, the activity check, the run check, the skip-in-use rule, the path matcher), and the test failed every time.
  • TestIdleRemoteSuspendHonoursReapBlockedIdle: with the setting disabled, side processes survive.
  • TestRemoteIdleSuspendIsScopedToTheRun: a window naming another run, with the stored run unchanged, is spared on the first sweep and ended (task suspended) on the next. When the stored run has moved on, it is left alone and the task is not re-asked at once. Removing the older flag in the awk program, or inverting the re-read, fails it.
  • TestReplacingARemoteTaskForgetsItsRemoteSession: after ClearTaskPlacement, a remote task has no session ID and HasLocalState no longer pins it here; a local task keeps its ID. Fails without the fix.
  • TestHostSweepFailsWithoutPs: the host script exits non-zero and says why.
  • TestRetryOfSuspendedRemoteTaskResumesItsSession: covers the recorded session, the newest-session fallback (a recorded but deleted ID, and subagent transcripts, are skipped), the --resume launch line, and the --session-id launch line.
  • TestRetryOfRemoteTaskBlocksWhenTheSessionCannotBeLookedUp: an unreachable host, and a missing worktree, fail with the message and never start an agent.
  • golangci-lint v2.13.2 (CI's version): 0 issues.
  • go test -race -p 4 ./... on macOS: everything passes except tests that also fail on origin/main here:
    • the /private/var executor tests;
    • reaper TestLivePanePIDs_HonorsAgentSocket;
    • ui TestRemoteAttachDropsALeftOverViewPairing and TestDetectMimeType…/typescript;
    • TestEnsureTaskWindow_ExportedEnvVarsReachLiveShellPane, which uses real tmux and flaked here (passes 2 of 3 runs alone).
  • The Linux /proc/<pid>/cwd path was not run locally (no Linux box or container here). CI's ubuntu job runs it.
  • Not run against a real host's sessions, and the daemon was not restarted.

Notes / follow-ups

  • Off Linux, a sidekiq whose command line doesn't name the worktree is not found, since there is no /proc cwd to read.
  • A resumed remote run gets only the reply, as a local resume does, not the completion-signal/MCP-proxy instructions again.
  • I assume claude --resume <id> keeps appending to the same session ID. If a Claude version forks a new ID on resume, a second resume would reopen the recorded (older) transcript rather than the newest.
  • A task closed before reap_blocked_idle passes has its windows ended by fix: end a placed task's remote agent session when it is done or archived #823's path, but its side processes are not reaped. That is unchanged from before.

🤖 Generated with Claude Code

bborn and others added 5 commits October 8, 2026 09:07
suspendIdleBlockedTasks finds an agent through this machine's tmux and ps,
so a task placed on another host was never suspended, and neither were the
dev servers it started there. Forty such tasks filled a staging host's
memory and swap for twenty days.

The per-host sweep from #823 now also ends the task-<id>/-shell windows of
placed tasks blocked longer than idle_suspend_timeout (in this coordinator's
session only), and, once reap_blocked_idle is also reached, their side
processes, matched by worktree path or puma/sidekiq title with the same
anchoring ty sessions cleanup uses. One ssh call per host, same backoff;
nothing is deleted.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… it is retried

A retry of a remotely placed task started a fresh Claude with the reply
appended to its prompt, so replying to a task the idle sweep had suspended
there would have lost the conversation. runRemoteSession now looks for the
task's newest session in its worktree's Claude project dir on the host and
launches claude --resume <id> with the reply, as a local retry does, falling
back to a fresh start when there is none.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review of #828: on a fleet host ty logs in as the user that also runs the
host's own TaskYou and other coordinators' tasks, so the sweep must not
reach anything but the idle task's own leftovers.

- Side processes: only this user's (ps -u), never an agent or tmux, never
  anything with a live tmux pane as itself or an ancestor (panes on every
  tmux socket of the user), never the sweep itself. On Linux matched by
  /proc/<pid>/cwd inside the recorded worktree (finds sidekiq, foreman,
  bin/dev too); elsewhere by the full worktree path or a [<dir>] puma title.
- No <id>- fallback: a task with no recorded worktree has its windows ended
  and its side processes left alone, and the log says so.
- Idle by the pane too: a window with activity inside the timeout leaves the
  task alone (windows and side processes), and it is not asked about again
  until it could be idle.
- Scoped to the run: an agent window whose start command names another
  WORKTREE_RUN_ID is a newer run's and is left alone.
- Window names are always sent; DaemonSession only decides whether to ask.
- A host without awk or ps is an error (retried), not 'nothing to end'.
- Resume: a session lookup that cannot reach the host blocks the retry
  instead of starting a fresh conversation. Placed Claude runs now start
  with --session-id <uuid>, recorded, and a retry resumes that exact
  session when the host still has it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ClearTaskPlacement (ty retry --replace) now forgets the Claude session ID
  of a task that was placed on another host. The session is on that host;
  left in the row, HasLocalState read it as a local first attempt and
  pinned the task to this machine without asking the resolver. A local
  placement's session is kept.
- A task-<id> window naming another run is spared only while it could be a
  newer run's. The sweep re-reads the stored run: if it is still the
  sweep's, the window is an older run's, and the next sweep ends it and
  suspends the task.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bborn
bborn merged commit 3a29064 into main Oct 8, 2026
4 checks passed
@bborn
bborn deleted the task/5717-idle-suspend-blocked-tasks-on-remote-hos branch October 8, 2026 15:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant