Skip to content
View kadirnar's full-sized avatar
🔥
Working from home
🔥
Working from home

Organizations

@goksenin-uav

Block or report kadirnar

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
kadirnar/README.md

Kadir Nar — Text-to-Speech & Audio. Cream mochi mascot wearing lavender headphones.

AI Research Engineer

Text-to-speech models, audio codecs,
training, inference, and evaluation.

text-to-speech  ·  audio codecs  ·  open source

🎙️ TTS models & training

  • VyvoTTS — LLM-based TTS training with SNAC and Mimi audio tokens.

  • Nar TTS — Speech-token TTS using causal language models and the Mimi audio codec.

  • Echo DACVAE — Experimental TTS with EchoDiT, flow matching, and DACVAE audio latents.

  • NAR-VAE — Experimental non-autoregressive TTS with flow matching and DACVAE latents.

🎧 Speech tools & agents

  • VoiceHub — Unified inference for speech synthesis, recognition, and voice activity detection.

  • VoicePlus — Native PyTorch inference pipelines for TTS, ASR, VAD, and audio codecs.

  • WhisperPlus — Whisper transcription, speaker diarization, summarization, and video captioning.

  • Voice Agent Next — Python runtime for speech-to-speech models and STT–LLM–TTS pipelines.

⚡ Audio codec optimization

  • Fast-DACVAE — DACVAE inference optimization with PyTorch compilation and CUDA graphs.

  • Fast-Mimi — Mimi codec inference optimization with fused kernels and CUDA graphs.

  • Fast-MOSS-Tokenizer — MOSS Audio Tokenizer v2 inference with CUDA Graph streaming and fused kernels.

🧪 Evaluation

  • VoiceHub Arena — TTS evaluation with intelligibility metrics, performance reports, and audio comparisons.

🌱 Earlier packages & open-source contributions

Packages

Contributions

Pinned Loading

  1. Vyvo-Labs/VyvoTTS Vyvo-Labs/VyvoTTS Public

    VyvoTTS: LLM-Based Text-to-Speech Training Framework

    Jupyter Notebook 262 40

  2. voicehub voicehub Public

    VoiceHub: A Unified Inference Interface for TTS Models

    Python 119 13

  3. whisper-plus whisper-plus Public

    WhisperPlus: Faster, Smarter, and More Capable 🚀

    Python 2k 147

  4. fast-dacvae fast-dacvae Public

    Fast inference engine for DACVAE, a neural audio codec that compresses and reconstructs audio using a convolutional encoder-decoder with a VAE bottleneck.

    Python 21 4

  5. nar-tts nar-tts Public

    Nar TTS: high-throughput Qwen3 + Mimi text-to-speech

    Python 22 1

  6. nar-vae nar-vae Public

    Low-cost non-autoregressive diffusion and flow-matching text-to-speech toolkit

    Python 9 2