NVIDIA's open 7B full-duplex speech-to-speech model: it listens while it talks, takes a voice and a role prompt, and answers in about 200 ms. English only, no tool calling.
Tencent's open-source 1.5B speech model: text to speech, voice cloning, and editing of recorded speech (words, pitch, speed, emotion, noise) from plain-language instructions.
NVIDIA's open-weight 100M-parameter model that tells who spoke when, for up to 8 speakers, live or recorded. Ranked first on VoiceArena's diarization benchmark at launch.