Phonon-2

Fermion Research's open English speech-to-text model: Parakeet-level accuracy in a 164 MB file, 174 times realtime on a MacBook Air, CC-BY-4.0.

Phonon-2 turns English speech into punctuated text, and it does it from a 164 MB file. That's about 15 times smaller than the NVIDIA model it was trained from, with nearly the same accuracy. On a MacBook Air it transcribes an hour of audio in about 20 seconds.

It comes from Fermion Research, a lab that builds very compact models: its Neutrino-1 language models use a ternary format, and Phonon-2 stores its encoder weights in about 2 bits each. The weights are open under CC-BY-4.0.

Phonon-2 launch banner: the Fermion Research spiral logo above the name Phonon-2 in a black serif font on an off-white background

Maker Fermion Research
Base model Distilled from NVIDIA's Parakeet TDT 0.6B v3, its "teacher"
Size 0.6 billion parameters, 164 MB download, 177 MB on disk
Language English only
Output Punctuated, capitalised text from 16 kHz audio
Accuracy 5.21% average word error on the 7 Open ASR Leaderboard sets
Licence CC-BY-4.0 for the weights, Apache 2.0 for the command line
Runs on Apple Silicon (MLX), Linux and Windows CPUs, NVIDIA GPUs

Accuracy and size

The comparison that matters is accuracy against download size. Lower word error rate (WER) is better.

Model Download Average WER
NVIDIA Parakeet TDT 0.6B v3, the teacher 2,508 MB 4.96%
Phonon-2 164 MB 5.21%
Moondream Parakeet Redux 178 MB 5.69%
NVIDIA Canary 180M Flash 737 MB 5.69%
Mistral Voxtral Mini 4B Realtime about 8,000 MB 6.12%
Phonon-1, the previous version 415 MB 6.56%
OpenAI Whisper large-v3-turbo 1,618 MB 6.58%

Fermion's claim: Phonon-2 is the most accurate open model under 900 MB, and every open model that scores better is at least 5.8 times its size. On meeting audio (AMI) and European Parliament speech (VoxPopuli), it even beats its teacher. It's weakest on earnings calls, where it trails the teacher by about 1 point.

Speed

Where Times realtime
MacBook Air M5, GPU (MLX) 174x
MacBook Air M5, CPU only 40x
Linux, 8 AMD Zen 5 cores 143x
Windows, 8 vCPU 21x
NVIDIA H100, batch of 128 6,680x

On the same MacBook Air, Fermion measured Parakeet TDT 0.6B v3 in FluidAudio's Core ML runtime at 105x, and whisper.cpp large-v3-turbo at 17x.

For business people

What it does: turns recordings and live speech into text, on your own device. No audio leaves the machine, and there's no per-minute bill.

Why it matters: good speech recognition used to mean a cloud API or a model of several gigabytes. At 164 MB, it fits on a laptop, a small server or inside an app download, and still runs faster than realtime on a plain CPU.

Where it fits: dictation, meeting and call transcripts, searchable archives of recorded audio, and subtitles for English content.

Costs: free. The licence allows commercial use with attribution.

The limits:

  • English only. The teacher speaks 25 European languages. Phonon-2 doesn't.
  • All benchmark and speed numbers are Fermion's own, scored by Fermion with the leaderboard's own code, not published by the leaderboard itself.
  • A young lab. Phonon-1 came out on 28th August 2026, and Phonon-2 in late September.

For technical people

  • Method: a distillation of Parakeet TDT 0.6B v3 with quantisation-aware training. Every encoder weight is 1 of 5 learned levels, about 2.1 bits per stored weight, with 6-bit tables elsewhere.
  • Fast path and exact path: the default decode on the Mac runs at 174x. An exact decode runs at 109x with the same 5.21% WER.
  • Packaging: 1 file for every platform, with a Core ML runtime for Apple devices announced but not yet out.

Install and run on Apple Silicon:

pip install fermion-research
pip install mlx mlx-audio mlx-lm soundfile scipy zstandard

fermion transcribe recording.wav --model phonon-2
fermion listen --model phonon-2    # live microphone transcription
fermion serve --model phonon-2     # local OpenAI-compatible transcription endpoint

On Linux or Windows, use Docker:

docker run --rm -v "$PWD":/audio ghcr.io/fermionresearch/phonon-cpu:2.0.2 \
  transcribe /audio/recording.wav --model phonon-2

The serve mode matters most: any tool that already calls OpenAI's transcription API can point at the local endpoint instead.

Detta, the app on top

Fermion also makes Detta, a Mac dictation app that runs Phonon-2 on the device: hold a key in any app, speak, and the text appears in the field.

The Detta window on macOS: a sidebar with Home, History, Transcribe and Dictionary, a list of dated dictations with speaker labels on some, and a stats card showing 18K total words and 161 words per minute

Value

My own dictation app, Dictee, runs the teacher model, Parakeet TDT 0.6B v3, on MLX (see Dictee: my own local dictation app ). Phonon-2 gives nearly the same accuracy at a tenth of the size and a faster runtime. The catch is language: I dictate in English, French and German, and Parakeet covers all 3. Phonon-2 would only replace it for English.

Where it fits for me today is batch work in English: call recordings, webinars and meeting audio, where 174x realtime turns an hour into 20 seconds on the laptop. The fermion serve endpoint makes it a drop-in test for any of my scripts that call a transcription API.

Further Reading

NicAI
Written by NicAI, Nic's AI assistant, for his personal knowledge base. Researched and drafted by the model, not hand-written by Nic. Verify anything you plan to act on.