Text-to-speech models, audio codecs,
training, inference, and evaluation.
text-to-speech · audio codecs · open source
-
VyvoTTS — LLM-based TTS training with SNAC and Mimi audio tokens.
-
Nar TTS — Speech-token TTS using causal language models and the Mimi audio codec.
-
Echo DACVAE — Experimental TTS with EchoDiT, flow matching, and DACVAE audio latents.
-
NAR-VAE — Experimental non-autoregressive TTS with flow matching and DACVAE latents.
-
VoiceHub — Unified inference for speech synthesis, recognition, and voice activity detection.
-
VoicePlus — Native PyTorch inference pipelines for TTS, ASR, VAD, and audio codecs.
-
WhisperPlus — Whisper transcription, speaker diarization, summarization, and video captioning.
-
Voice Agent Next — Python runtime for speech-to-speech models and STT–LLM–TTS pipelines.
-
Fast-DACVAE — DACVAE inference optimization with PyTorch compilation and CUDA graphs.
-
Fast-Mimi — Mimi codec inference optimization with fused kernels and CUDA graphs.
-
Fast-MOSS-Tokenizer — MOSS Audio Tokenizer v2 inference with CUDA Graph streaming and fused kernels.
- VoiceHub Arena — TTS evaluation with intelligibility metrics, performance reports, and audio comparisons.
- Tracking: Norfair · OC-SORT · ByteTrack · StrongSORT · SORT
- Detection & segmentation: MetaSeg · YOLOv7 · YOLOv6 · YOLOX
- Image restoration: CodeFormer · BSRGAN
- Speech & generative AI: Insanely Fast Whisper · Diffusers #3590 · Diffusers #3586 · ComfyUI
- ML tools: Evaluate · PyTorch Examples · LlamaIndex · AutoLLM
- Computer vision: SAHI #486 · SAHI #322 · SAHI #501 · Norfair · StrongSORT · YOLOv7 · YOLOv6 · Kornia





