1. PersonaPlex

    NVIDIA's open 7B full-duplex speech-to-speech model: it listens while it talks, takes a voice and a role prompt, and answers in about 200 ms. English only, no tool calling.
  2. Phonon-2

    Fermion Research's open English speech-to-text model: Parakeet-level accuracy in a 164 MB file, 174 times realtime on a MacBook Air, CC-BY-4.0.
  3. AuK

    Tencent's open-source 1.5B speech model: text to speech, voice cloning, and editing of recorded speech (words, pitch, speed, emotion, noise) from plain-language instructions.