Skip to content

Windows file:// fix, OmniVoice tail/tensor fixes, Omni Flash via worker + provider pacing - #2

Open
khnhlt wants to merge 6 commits into
mainfrom
fix/windows-file-url-paths
Open

khnhlt wants to merge 6 commits into
mainfrom
fix/windows-file-url-paths

Conversation

@khnhlt

@khnhlt khnhlt commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Summary

Six commits from producing a real 8-scene, 62s video ("Người gieo hạt") end-to-end on Windows with Omni Flash + OmniVoice. Each commit fixes one thing found by running the pipeline, and each ships with a regression test (RED → GREEN). Suite: 454 passed.

Commit d59257e is already proposed upstream as crisng95#69; the other five are new.

Commits

Commit What / why
fix d59257e file:// URLs mis-parsed on Windows (urlparse puts C: in netloc) → every local clip looked missing. New file_url_to_path() used at 5 call sites.
fix 5e5dec2 OmniVoice returns 1-D numpy; torchaudio.save needs 2-D → wrapped. Also honour TTS_DEVICE instead of hard-coded cpu.
feat 328df2a project.video_model_family (veo | omni_flash, default veo, schema migration). Worker was hard-wired to Veo; Omni Flash was reachable only via the ad-hoc /api/flow/generate-video, which doesn't write back to the scene. Now _run_video dispatches on the family and review/concat/retry work unchanged. Also adds the column to the crud._update whitelist (PATCH was a silent no-op otherwise — covered by test).
feat 0b790bc FLOW_MAX_CONCURRENT / FLOW_COOLDOWN_S env. ~30 generates 10s apart from a background tab degraded the reCAPTCHA session score until Flow returned PUBLIC_ERROR_UNUSUAL_ACTIVITY on every call, while the same account still generated fine by hand. Defaults unchanged.
fix cb1fd33 OmniVoice stops at the last voiced sample (no release tail) → final vowel clipped, and "trim video to narrator length" cuts mid-sound. Fade last 60 ms + pad 0.4 s at save time.
feat 2f2c5d8 scene.duration → Omni Flash clip length, rounded up to 4/6/8/10 s. A 7.6 s narration can't breathe in an 8 s clip; the only post fix is a frozen frame. Veo ignores the field.

Test plan

  • pytest in agent venv: 454 passed, 4 skipped (Windows, ffmpeg on PATH)
  • pytest tests/unit/test_tts_scripts.py in TTS venv (torch/omnivoice): 5 passed
  • Live: 8 scenes generated through the worker with video_model_family=omni_flash → ffprobe 1280×720 / 8 s / AAC (Omni signature; Veo would be 1080p); one scene regenerated at duration=10 → 10.005 s
  • Live: /videos/{vid}/narrate with the patched TTS → wavs end with 0.4 s silence, no clipped vowel
  • Live: after FLOW_COOLDOWN_S=45 + tab reload, UNUSUAL_ACTIVITY cleared and generation resumed

Notes for review

Khanh added 6 commits September 29, 2026 19:44
The pipeline stores local media as "file://" + str(path). On Windows that
is file://C:\...\x.mp4, which urlparse splits into netloc="C:\..." and an
empty path, so every Path(parsed.path) caller decided the clip was missing
and fell through to an HTTP download that then failed.

Add agent.utils.paths.file_url_to_path (netloc+path via url2pathname, so
both as_uri() and the naive form work) and use it in video_reviewer,
api/videos, api/projects and assistant_provider. Make the two
test_cli_providers assertions spell /tmp the way the host OS does.

Windows: 12 failed -> 0 failed (445 passed).
OmniVoice.generate() returns list[numpy.ndarray] (1-D). torchaudio.save
needs (channels, samples), so every TTS call failed at the last step with
'Expected 2D Tensor, got 1D' after the model had generated fine. Wrap the
item via a shared _to_wav_tensor helper in both inline scripts.

Also pass config.TTS_DEVICE through to device_map instead of hardcoding
cpu — TTS_DEVICE already existed in config but was never used. On an RTX
3060 a template takes ~23s including model load vs minutes on CPU.
The queue path (GENERATE_VIDEO -> FlowProvider._run_video) was hardwired
to Veo. Omni Flash was reachable only through the ad-hoc
/api/flow/generate-video endpoint, which does not persist results to the
scene, so choosing Omni meant leaving the pipeline (no review, no concat,
no retry-repoll).

Add project.video_model_family ('veo' | 'omni_flash', default veo, with
schema migration), thread it through ProjectCreate/ProjectUpdate, the SDK
model/repository, and into job.extra. _run_video dispatches on it:
omni_flash -> generate_omni_flash_first_frame_video / _first_last_video,
which return the same batch-operation shape, so polling and the
already-submitted guard are shared. Defaults come from
OMNI_FLASH_DURATION_S / OMNI_FLASH_RESOLUTION.
Each generate mints a reCAPTCHA in the Flow tab. A worker firing ~30 of
them 10s apart from a background tab degrades the session score until
Flow returns PUBLIC_ERROR_UNUSUAL_ACTIVITY on every call, while the same
account still generates fine by hand. Expose the provider's max_concurrent
and cooldown_s via env so operators can slow down without editing code;
defaults unchanged.
…id-vowel

OmniVoice returns audio that stops at the last voiced sample: zero release
tail, final vowel clipped. Any pipeline that trims video to narrator length
then cuts exactly where the voice is still ringing, which reads as an abrupt
scene change. Fade the last 60 ms and append 0.4 s of silence inside the
save path so every wav ends cleanly and downstream trims get a breath.
A scene whose narration runs 7.6s cannot breathe inside an 8s clip; the
only post fix is a frozen last frame, which reads as a glitch. Pass
scene.duration into job.extra and round it UP to the next supported Omni
Flash step (4/6/8/10) in the provider — rounding down would clip the
voice. Veo ignores the field. Defaults unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant