
PersonaPlex is an open-weight voice model from NVIDIA's Applied Deep Learning Research (ADLR) team. It is full duplex: 1 model hears the user and speaks at the same time, so it handles interruptions, pauses and "mm-hm" backchannels the way a person does. Most voice agents chain 3 models (speech-to-text, an LLM, text-to-speech) and take turns like a walkie-talkie. PersonaPlex skips the chain.
What it adds to the open model it builds on, Kyutai's Moshi, is control. You give it 2 prompts before the conversation starts:
- a voice prompt, a short audio sample that sets the voice, accent and speaking style
- a text prompt, which sets the role, the name, the company and the facts it should know
So you can have a 200 ms-latency receptionist called Owen who knows the menu prices, in a voice you choose. It is English only, and it can't call tools.
Key facts
| Maker | NVIDIA ADLR (Rajarshi Roy, Bryan Catanzaro and others) |
| Released | 15th January 2026. Paper on arXiv in February 2026 |
| Model | personaplex-7b-v1, 7B parameters, fine-tuned from Kyutai's Moshi |
| Language | Python and PyTorch. Model in English only |
| Audio | 24 kHz in and out, speech plus a parallel text stream |
| Licence | Code MIT. Weights under the NVIDIA Open Model License, commercial use allowed |
| Hardware | Tested on NVIDIA A100 80 GB. Officially Ampere and Hopper GPUs on Linux |
| Traction | About 10,700 GitHub stars, 166,000 Hugging Face downloads in the last month |
| Maintenance | Last code change 23rd January 2026. A research release, not an active product |
What it can do
- Natural turn-taking. It answers in about 0.2 s, stops when you cut in, waits through your pauses and says "yeah" or "oh okay" while you talk.
- Voice control. 18 preset voices ship with the repo: 8 "natural" (
NATF0-NATF3,NATM0-NATM3) and 10 "variety" (VARF0-VARF4,VARM0-VARM4). - Role control. A plain-text prompt makes it a teacher, a bank agent, a medical receptionist or a restaurant order-taker.
- Some generalisation. Because Moshi sits on Kyutai's Helium LLM, it will play roles it wasn't trained on. NVIDIA's demo has it as an astronaut handling a reactor meltdown on the way to Mars.

For business people
- The problem it solves: voice bots feel robotic because they wait for silence, then think, then talk. PersonaPlex talks like a person: fast, interruptible, with small acknowledgements. That is most of what makes a call feel human.
- Where it fits: first-line customer service, appointment booking, intake calls, tutoring, game characters, and the voice behind a conversational avatar.
- Cost shape: the model is free, and NVIDIA allows commercial use. You pay for the GPU. NVIDIA tested on an A100 80 GB: plan for data-centre hardware, not a laptop.
- The limits are real:
- No tool calling. It can't look up an order, check a calendar or book a slot. Everything it knows must be in the text prompt.
- English only. For German or French callers it is not an option.
- Small brain. The 7B backbone is good at conversation and weak at reasoning. Users on GitHub report hallucinations and off-script answers.
- No support. NVIDIA hasn't touched the code since January 2026.
- When to choose it: for a demo, a prototype or research on natural turn-taking, it is the best open option with a commercial licence. For a production agent that must act on systems, a 3-model pipeline or a hosted service such as Fonio is still the practical choice (see Fonio).
For technical people
How it works
PersonaPlex keeps Moshi's architecture and changes the training:
- Mimi codec: turns 24 kHz audio into discrete tokens and back.
- Temporal transformer: the 7B backbone. At each audio frame it models 3 parallel streams: the user's audio, the agent's text and the agent's audio. Because the user's stream never stops, the model "hears" while it speaks.
- Depth transformer: predicts the stack of codec tokens for the agent's audio in each frame.
- Hybrid system prompt: before generation, the model receives the voice prompt as agent audio, a pause, the text prompt as agent text, and another pause. That prefix is what Moshi didn't have.
Training data
| Source | Conversations | Hours |
|---|---|---|
| Real calls (Fisher English) | 7,303 | 1,217 |
| Synthetic assistant | 39,322 | 410 |
| Synthetic customer service | 105,410 | 1,840 |
LLMs (Qwen3-32B and GPT-OSS-120B) wrote the synthetic transcripts, and Chatterbox TTS voiced them. GPT-OSS-120B also back-labelled the real Fisher calls with prompts. The team's finding: under 5,000 hours of directed data was enough to teach a pretrained duplex model to follow a role.
Benchmarks
NVIDIA's own numbers on FullDuplexBench and its in-house ServiceDuplexBench:
| Model | Turn-taking score (avg) | Latency (avg) | Task adherence (0 to 5) |
|---|---|---|---|
| PersonaPlex | 82.1 | 0.205 s | 4.34 |
| Gemini Live (2.0) | 75.5 | 1.242 s | 4.05 |
| Moshi | 65.3 | 0.261 s | 1.26 |
| Freeze Omni | 54.7 | 1.181 s | 3.82 |
PersonaPlex wins on the averages, not on every metric. Moshi interrupts and takes turns slightly better but fails pause handling (1.8 out of 100). Gemini Live scores higher on the customer service tasks. Qwen 2.5 Omni beats it on FullDuplexBench task adherence.
Install and run
- Install the Opus codec library:
sudo apt install libopus-dev(ordnf install opus-devel). - Clone the repo and install the bundled Moshi fork.
- Accept the model licence on Hugging Face and export a token.
- Start the server and open the web UI on port 8998.
git clone https://github.com/NVIDIA/personaplex && cd personaplex
pip install moshi/.
export HF_TOKEN=<your token>
# Live web UI with temporary SSL certs
SSL_DIR=$(mktemp -d); python -m moshi.server --ssl "$SSL_DIR"
# Not enough VRAM: offload layers to the CPU (needs accelerate)
SSL_DIR=$(mktemp -d); python -m moshi.server --ssl "$SSL_DIR" --cpu-offload
Offline, it streams a WAV file in and writes the reply as WAV plus a JSON transcript:
python -m moshi.offline \
--voice-prompt "NATM1.pt" \
--text-prompt "You work for Jerusalem Shakshuka which is a restaurant and your name is Owen Foster. Information: ..." \
--input-wav input.wav \
--output-wav output.wav --output-text output.json
The text prompt format matters. The model was trained on You work for {company} which is a {type} and your name is {name}. Information: {facts} for service roles, and on You enjoy having a good conversation. plus a topic for casual talk.
On a Mac
The official code needs CUDA. 2 community ports run it on Apple Silicon with MLX:
- speech-swift by Ivan (soniqo): native Swift, the checkpoint quantised from 16.7 GB to 5.3 GB at 4 bits, about 68 ms per step, slightly faster than real time. Developer benchmarks, not independent ones.
- personaplex-mlx: a Python MLX port with local and web modes.
Gotchas
- VRAM. The bf16 checkpoint is about 16.7 GB. On smaller cards use
--cpu-offloador a community 4-bit or 8-bit build, and expect slower answers. - Blackwell GPUs need the CUDA 13 PyTorch wheels:
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130. - DGX Spark produces choppy, unusable audio. It is the most-discussed open issue, with 58 comments.
- Browser audio needs HTTPS for the microphone, which is why the server makes temporary SSL certificates.
- Reproducible runs. Pass
--seedfor repeatable offline output, as the README examples do.
PersonaPlex or VoiceChat
In September 2026 NVIDIA released NemotronLabs VoiceChat 11B, a second full-duplex model built on a different stack. The 2 trade off in opposite directions:
| PersonaPlex 7B | VoiceChat 11B | |
|---|---|---|
| Released | January 2026 | September 2026 |
| Backbone | Moshi and Helium | Nemotron Nano V2 9B |
| Tool calling | No | Yes, with spoken "on hold" lines |
| Latency | About 0.2 s | About 0.45 s |
| Licence | Commercial use allowed | Research only |
Pros and cons
Pros
- The best turn-taking scores among the open models NVIDIA tested: about 0.2 s, with real pause handling.
- Voice and role set by a prompt, no fine-tuning needed.
- Commercial-use licence on the weights, MIT on the code.
- Runs locally, on a GPU or on a Mac through the MLX ports.
Cons
- English only.
- No tools, no retrieval: the prompt is the whole knowledge base.
- 7B conversational backbone, so weak reasoning and some hallucination.
- Frozen research release since January 2026.
- Needs a serious NVIDIA GPU for the official code.
My take
PersonaPlex is the best reference I know for how a voice agent should feel. The 0.2 s reply and the small "yeah" while you talk do more for perceived quality than a smarter LLM behind a slow pipeline. That is the bar any conversational avatar has to clear, including the ones I sell at Kaltura (background in Research: Conversational Interfaces and AI Video Agents (Digital Humans)).
As a product component it falls short for me today. No German, no tool calling, so it can't drive Home Assistant or book anything, and NVIDIA has moved on to VoiceChat. The thing I'd do is run the speech-swift port on the Mac Studio for an evening, to hear full duplex locally and calibrate my ear before judging vendor demos.
Further reading



- Moshi, the base model, by Kyutai
- FullDuplexBench, the benchmark behind the turn-taking scores
- PersonaPlex 7B on Apple Silicon in native Swift with MLX (Ivan, 23rd February 2026)
- NVIDIA Open Model License
- NVIDIA Nemotron 3 Diarization, another open NVIDIA speech model