Skip to content

Fix TTS voice being ignored/inconsistent in read-aloud and video export - #189

Merged
davior merged 1 commit into
mainfrom
claude/tender-fermat-qy5vuq
Sep 12, 2026
Merged

davior merged 1 commit into
mainfrom
claude/tender-fermat-qy5vuq

Conversation

@davior

@davior davior commented Sep 12, 2026

Copy link
Copy Markdown
Owner

Summary

Fixes the reported bug: setting the read-aloud voice to "Rex" (a fal.ai voice) did not take effect in AI Assistant read-aloud or in MP4 export, and some rendered videos changed voice partway through.

Three separate root causes:

  • AI Assistant read-aloud / voice mode: AIConversationPanel.tsx and useVoiceMode.ts both called useTextToSpeech() with no model option. When the fal.ai fallback path is used, the frontend defaults an unset voice to the hardcoded DEFAULT_TTS_VOICE ("Aria") instead of the user's configured voice. Fixed by passing the settings store's current voice through, matching how EditorView.tsx and SpeechSettings.tsx already do it.
  • Article-to-video export: the render worker's _tts_caller never passed the user's selected voice into synthesize_tts_bytes at all, so the fal.ai path always fell back to the current TTS model's first listed voice — never the user's actual choice. Added load_selected_voice() to read the per-user voice setting and pass it through for every render.
  • Voice changing mid-video: when the read-aloud provider is "auto" and both a Deepgram and fal.ai key are configured, provider selection is re-resolved on every synthesis call, and a transient Deepgram failure silently falls back to fal.ai for just that one call. For a narration job that makes one call per text chunk, this could swap a single chunk to a different engine/voice mid-render. The provider is now resolved once per render job and pinned for the whole render, so a mid-job hiccup now fails the render (which is safe to retry) instead of producing a video with an inconsistent voice.

Test plan

  • Verified the fal.ai voice-resolution path (_fal_tts, synthesize_tts_bytes) and confirmed the new voice/provider_override plumbing is used consistently by the render worker.
  • Reviewed the frontend hooks (useTextToSpeech, useVoiceMode, AIConversationPanel) to confirm the selected voice is now threaded through everywhere useTextToSpeech is instantiated.
  • Manual verification in a running instance (set voice to a non-default fal.ai voice, confirm it's used in AI Assistant read-aloud and in an exported MP4) was not possible in this sandbox — dependencies (fal_client, pydantic, frontend node_modules) aren't installed, so this couldn't be exercised end-to-end. Please verify manually before merging.

🤖 Generated with Claude Code

https://claude.ai/code/session_01R54PjuaaN6QnvxVCcxiuVs


Generated by Claude Code

Three related bugs caused a chosen fal.ai voice (e.g. "Rex") to be
silently replaced:

- AI Assistant read-aloud and voice mode called useTextToSpeech()
  without a voice, so any fal.ai fallback used the hardcoded default
  voice instead of the one configured in Settings -> Speech.
- The article-to-video renderer never forwarded the selected voice at
  all, so every export used the current model's first listed voice.
- Article-to-video narration re-resolved the "auto" Deepgram/fal.ai
  provider choice per chunk, so a single transient Deepgram failure
  mid-render silently swapped just that chunk to fal.ai's voice,
  producing a video whose narrator changes partway through. The
  provider is now resolved once per render job and pinned for every
  chunk, so a failure now fails the render instead of mixing voices.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R54PjuaaN6QnvxVCcxiuVs
@davior
davior marked this pull request as ready for review September 12, 2026 13:50
@davior
davior merged commit fe888fc into main Sep 12, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants