Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
-
Updated
Sep 22, 2026 - Python
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
Turn detection for full-duplex dialogue communication
Real-time Turn-taking, Backchannel, and Head-nodding Prediction for Full-duplex Interaction
[EMNLP 2026 Main] Speaking While Listening: a survey and empirical audit of full-duplex spoken dialogue systems — L0–L3 architectural hierarchy, T×I×R interaction ontology, and a five-state decision machine, with a curated list of models, datasets, and benchmarks.
DuplexJev-32B and DuplexJev-4B: speech decision models for full-duplex voice agents. Typed single-token readout from a frozen speech LLM, served with vLLM. Try it: api.adventists.cn/duplexjev
GenPark AI Agent Skill - Conversational audio acoustic filler injector masking LLM inference and TTS latency with contextual conversational cues.
GenPark AI Agent Skill - Conversational audio acoustic filler injector masking LLM inference and TTS latency with contextual conversational cues.
GenPark AI Agent Skill - Phonetic Soundex and acoustic confusion matrix resolver correcting domain-specific STT transcription errors.
GenPark AI Agent Skill - Real-time conversational voice activity detection, dynamic silence endpointing and barge-in interruption arbitrator.
GenPark AI Agent Skill - PSTN/SIP voice telephony state machine managing DTMF tone decoders, call transfer handoffs and IVR navigation.
GenPark AI Agent Skill - Real-time conversational voice activity detection, dynamic silence endpointing and barge-in interruption arbitrator.
GenPark AI Agent Skill - PSTN/SIP voice telephony state machine managing DTMF tone decoders, call transfer handoffs and IVR navigation.
GenPark AI Agent Skill - Real-time RTP and WebSocket audio packet jitter buffer optimizer smoothing network latency and clock drift for speech streaming.
GenPark AI Agent Skill - Phonetic Soundex and acoustic confusion matrix resolver correcting domain-specific STT transcription errors.
GenPark AI Agent Skill - Real-time RTP and WebSocket audio packet jitter buffer optimizer smoothing network latency and clock drift for speech streaming.
Real-time acoustic VAD & semantic turn-completion endpoint detector distinguishing mid-sentence hesitation pauses from finished speech turns.
Low-latency streaming prosody, SSML, and emotion pacing modulator dynamically tailoring speech pitch, cadence, and pause inflections.
Packet ordering simulation and inter-arrival jitter telemetry.
End-to-end conversational voice agent latency telemetry profiler tracking VAD, STT, LLM-TTFT, TTS-TTFB, and playout bottlenecks.
Full-duplex conversational barge-in arbitrator classifying user interruptions, backchannels, and acoustic echo with context rollback.
To associate your repository with the turn-taking topic, visit your repo's landing page and select "manage topics."