Skip to content

[Bug]: Codex HTTP 429 retry exhaustion hides Limited state and manual Resume #16288

Description

@coygeek

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Impact

Minor bug or occasional failure. The original stop interrupted an active task; a new user message continued the same native session.

Workaround

The user typed go on, and a second run began in the same native provider session with further tool activity.

Summary

A Codex turn stopped with an explicit HTTP 429 after exhausting provider retries. T3 Code displayed Failed and a generic provider error instead of the documented Limited state and manual Resume control. This interrupted an active task. Sending a new user message later started a continuation in the same native provider session, so this report concerns the missing limit classification and recovery control, not an unrecoverable conversation.

Steps to reproduce

These are the observed incident conditions, not a deterministic reproduction. A fresh rate-limit reproduction was not run.

  1. Start a Codex coding task in the desktop app using the runtime listed below. The observed turn performed tool calls and delegated work to three subagents through a custom provider route.
  2. While the turn is active, the provider exhausts retries and returns exceeded retry limit, last status: 429 Too Many Requests with the error code responseTooManyFailedAttempts. The request conditions that caused the provider to return 429 are unavailable.
  3. Inspect the thread immediately after termination. The supplied screenshot shows Failed, a Provider error item, and an empty composer with the send control rather than the documented limited-turn Resume control.
  4. Inspect the persisted failure classification. The observed incident records provider_error, not usage_limit.
  5. As a separately observed workaround, send a new continuation message. A later user message started a second run using the same native provider session and further tool activity appeared.

Expected behavior

T3 Code's documentation defines Limited as a provider stop on a usage or rate limit and promises manual Resume in an empty composer. Providers without a reset time still offer manual retry. An explicit HTTP 429 retry-exhaustion stop should retain that rate-limit recovery behavior. This does not require scheduling an automatic continuation when no reset time is known, or treating every responseTooManyFailedAttempts error as a rate limit.

Actual behavior

The original run ended in failed. The persisted provider thread became idle, with no queued continuation or unfinished run at the initial inspection. The error item contained the following sanitized fields. Two subagents ended in failed and one ended in completed; their failure causes were not independently established.

{
  "class": "provider_error",
  "message": "exceeded retry limit, last status: 429 Too Many Requests",
  "code": "responseTooManyFailedAttempts",
  "retryable": null
}

The screenshot's sidebar showed Failed. A later user-authored continuation message started another run in the same native provider session. Manual continuation therefore worked in this incident, but the rate-limit-specific state and empty-composer recovery control were absent at the original stop. No controlled failure rerun or fix verification was performed.

Evidence

  • Expected source: The installed build's thread recovery documentation, under Inspect agent work, defines the Limited state, manual Resume, reset-time scheduling, and manual retry without a reset time.
  • Failure source: Read-only inspection of the live v2 store's orchestration_v2_projection_turn_items.payload_json.failure, run and provider-thread projections, and a supplied screenshot of the same failed thread. The sanitized failure fields are reproduced above. A subsequent snapshot records a user-authored continuation and a second run with further tool activity.
  • Evidence provenance: observed
  • Local verification: not-run
  • Reproduction completeness: incomplete
  • Missing fact: The provider-side request and rate-limit conditions, or an isolated fixture that faithfully produces this explicit HTTP 429 terminal error; original response headers and reset information are unavailable.

The persisted failure, run state, and later continuation were inspected directly; the UI evidence comes from the user-supplied screenshot. Investigation inspected retained incident evidence and source documentation and did not induce a new provider rate limit. The original screenshot and transcripts contain private context and are not attached. Inline evidence preserves the error and state needed for this report while omitting session identifiers, private task content, and machine paths. The persisted failure does not identify whether the 429 originated at the custom route or its upstream service.

Source inspection of the installed revision provides investigation locators. In apps/server/src/orchestration-v2/Adapters/CodexAdapterV2.ts, codexErrorInfoCode extracts the error variant key, and both terminal failure mappings classify only usageLimitExceeded and rateLimitExceeded as usage_limit. In apps/web/src/components/ChatView.tsx, resumableRunId admits a failed run only when the runtime and root failure both have usage_limit. These source observations explain the available recovery path; they do not establish the upstream cause of the 429 or replace a fresh reproduction.

Restoration check

With an isolated Codex turn that terminates in explicit HTTP 429 retry exhaustion, verify that the thread exposes the documented Limited state and manual Resume with an empty composer, including when there is no reset time. After the provider accepts requests again, manual Resume should continue the same conversation and preserve previous work. Use a non-429 retry-exhaustion error as a control: it should keep its applicable error handling rather than gain a misleading rate-limit state. These checks remain unexecuted in this report.

Version or commit

The running desktop installation reports T3 Code Nightly 0.0.46-nightly.20261005.2702 in both its app metadata and bundled package manifest. The retained native transcript reports Codex CLI 0.160.1; the selected model was gpt-6.1-sol, with ultra reasoning and the default service tier. The session used a custom provider route. Its upstream request path and original HTTP headers are unknown. The documentation citation pins the source revision reported by the bundled manifest. No reproduction on another build was attempted.

Triage assessment

  • Impact level: P2
  • Assessment status: supported
  • Impact basis: One retained incident stopped an active coding task and hid the documented rate-limit state and manual recovery control. Wider frequency and affected-provider reach are unknown; no data loss was observed.
  • Workaround status: available
  • Workaround basis: A later user-authored continuation message started a second run in the same native provider session with further tool activity. This preserves the conversation but requires the user to notice the stop and send another message; completion of the coding task was not evaluated.

Activity

  1. juliusmarminge commented on Oct 6, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks @coygeek for the thorough report and for pointing at the right files. The stored failure was enough to trace this. I didn't reproduce a fresh 429, but on current main (cfa4f765ec) the classifier explains the Failed state.

    What I found

    • Codex reports retry exhaustion as codexErrorInfo: { responseTooManyFailedAttempts: { httpStatusCode } }. codexErrorInfoCode keeps only the variant key, so the HTTP status never reaches the failure record. That matches the stored code: "responseTooManyFailedAttempts" and retryable: null.
    • makeRootTerminalEvent, and the non-retrying error notification that seeds latestProviderFailure, map only usageLimitExceeded and rateLimitExceeded to usage_limit, so this variant becomes provider_error. The in-progress retry path treats names starting with http or responseStream as transport_error, and this name matches neither. The classifier test only covers those string codes.
    • threadErrorSummary copies that class onto the thread. The sidebar and mobile thread list show Limited only for usage_limit, and web/desktop Resume requires usage_limit on both the thread and the root failure (the orchestrator rejects manualContinuationOfRunId otherwise). That accounts for the Failed state, the Provider error item, and the missing Resume control.
    • A normal follow-up takes a different path. Failures already set threadDisposition: "reusable", which is why go on continued in the same native session.
    • Claude already treats HTTP 429 as usage_limit, and the thread docs describe Limited plus manual Resume for a rate-limit stop even without a reset time. Auto-resume still requires usageLimitResetAt, so reclassifying wouldn't schedule a continuation by itself. Custom routes often have no rate-limit snapshot, which fits this incident.
    • This looks distinct from feat(v2): show provider limit stops as Limited #12677 and feat(v2): resume limited threads when usage resets #12686, which cover the dedicated limit codes and reset-time resume rather than this HTTP-status wrapper.

    Likely fix area

    One option is treating httpStatusCode === 429 on responseTooManyFailedAttempts as usage_limit in both the terminal mapping and the error notification, and keeping that field in codexErrorInfoCode. Other statuses on the same variant would stay provider_error, since the variant name alone doesn't mean a limit was hit. httpConnectionFailed and the responseStream* variants carry the same status field and could be considered alongside it.

    One caveat: the stored failure has no raw codexErrorInfo, so this incident's 429 is inferred from the message and the schema. If a build reports 429 only in the message without httpStatusCode, status-only classification would still miss it.

    A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions