Skip to content

fix(indexer): crash the process when the ETL indexer fails to start - #1042

Merged
rickyrombo merged 1 commit into
mainfrom
fix/core-indexer-errgroup-hang
Sep 22, 2026
Merged

rickyrombo merged 1 commit into
mainfrom
fix/core-indexer-errgroup-hang

Conversation

@rickyrombo

Copy link
Copy Markdown
Contributor

What happened

Prod's core indexer stopped indexing at block 32,826,125 and stayed there for 8 hours. The pod reported Running 1/1 the whole time, with every background job (PrunePlays, IndexChallenges, Trending, …) logging normally. Plays, follows, and every other on-chain write were invisible to the API for the duration — this surfaced as a user's plays missing from the database.

Timeline from the pod logs:

Time (UTC) Event
10:22:26 GKE node upgrade recreates every pod; creator-1/audiusd-0 created
10:22:33 api/core-indexer starts, begins resolving chain ID from rpc.audius.co
10:22:33–10:23:37 attempts 1–29: 503 Service Unavailable, then core service not ready
10:23:33 audiusd container actually starts
10:23:39 attempt 30 fails → error initializing chain ID after 30 attempts

The ETL's initializeChainID uses 30 attempts at a flat 2s delay — a fixed ~58s budget. audiusd took ~67s. It missed by about six seconds.

Why it never recovered

InitializeChainID runs in etl.Indexer.Run() before indexBlocks() is launched, so Run() returned the error immediately. main.go:58 panics on that error, k8s would have restarted the pod into a now-healthy audiusd, and this would have self-healed in a minute.

It never got there:

eg := errgroup.Group{}          // plain Group, no context
eg.Go(func() error { return ci.aggregatesCalculator.Start(ctx) })
eg.Go(func() error { return ci.etlIndexer.Run() })
return eg.Wait()

errgroup.Group.Wait() blocks until all goroutines return, and a plain Group has no context to cancel siblings with. AggregatesCalculator.Start only returns on ctx.Done(), and that ctx is cancelled only by SIGTERM. So the ETL's fatal error sat in the errgroup while Wait() blocked forever on a healthy infinite loop. Start never returned and the process never exited.

The fix

errgroup.WithContext, so the first error cancels the sibling and Wait() returns it. main.go then panics as it was always meant to and k8s restarts the pod.

Graceful shutdown is unchanged: a SIGTERM-cancelled parent still propagates to gCtx, Start returns context.Canceled, and main.go ignores it. The parity jobs move to gCtx for the same reason — they should stop when the process is on its way down.

Recovery is automatic

Resume is GetLatestIndexedBlock() + 1 from etl_blocks, so a restart replays the gap rather than skipping it. Confirmed on the restarted pod: block source: polling prefetcher start_height=32826126, exactly where it left off.

Testing

Added TestAggregatesCalculatorStartReturnsOnCancelledContext, which locks in the property the fix depends on: the sibling goroutine actually unblocks when its context is cancelled. If that loop ever stops honoring ctx, Wait() silently goes back to hanging forever.

The fix itself isn't unit-covered — etlIndexer is a concrete *etl.Indexer and aggregatesCalculator needs live DB pools, so neither member is injectable without a refactor. Happy to add that seam as a follow-up if it's wanted.

Follow-ups not in this PR

  • No probes or health server on core-indexer. The deployment has no liveness, readiness, or startup probe, and main.go's indexer case doesn't start a health server the way the solana-indexer and eth-indexer cases do. /health_check?max_core_indexer_block_diff=N already returns 500 on excess lag — nothing calls it. This is why a wedged process went unnoticed for 8 hours.
  • Chain-ID retry budget is shorter than a cold audiusd start. 58s flat makes any simultaneous restart a coin flip. Related: SetCheckReadiness(true) is set, but awaitReadiness runs after chain-ID init and only logs on failure, so it can't help here.

🤖 Generated with Claude Code

CoreIndexer.Start used a plain errgroup.Group, so an error from
etlIndexer.Run() was recorded but never propagated: Wait() blocks until
every goroutine returns, and AggregatesCalculator.Start only returns on
ctx.Done(). With no context linking the two, a fatal ETL error left the
process alive with block indexing permanently dead.

This bit prod today. A GKE node upgrade restarted every pod at once;
audiusd took ~67s to become ready, longer than the ETL's fixed 30x2s
chain-ID retry budget, so Run() returned "initialize chain ID after 30
attempts" before indexBlocks() ever started. main.go panics on that
error and k8s would have restarted the pod into a healthy audiusd, but
Start never returned. The pod reported Running 1/1 for 8 hours while
every other background job logged normally and the indexer sat 25k
blocks behind, silently dropping plays, follows and every other
on-chain write from the API's view.

errgroup.WithContext cancels the sibling on the first error, so
Wait() now returns the ETL error and main.go panics as intended.
Graceful shutdown is unchanged: a SIGTERM-cancelled parent still
surfaces as context.Canceled, which main.go ignores.

The parity jobs move to the group context for the same reason — they
should stop when the process is on its way down.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@rickyrombo
rickyrombo merged commit 5708b57 into main Sep 22, 2026
2 checks passed
@rickyrombo
rickyrombo deleted the fix/core-indexer-errgroup-hang branch September 22, 2026 19:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant