AI Model Release Tracker

Open speech models add real-time diarization, TTS, and voice-agent benchmarks

Open speech models add real-time diarization, TTS, and voice-agent benchmarks

NVIDIA's Nemotron 3 Diarization stands out as a 100M-parameter open-weight, commercially licensed model with NeMo code, Hugging Face weights, demos, eight-speaker overlap tracking, and detailed streaming DER/latency results. Google DeepMind's Gemini 3.8 Flash TTS models and Guava's Daytona voice model add expressive multilingual synthesis, safeguards, and an open voice-agent benchmark, though independent validation remains limited.

Sources (2)
Updated Sep 23, 2026