You have 100 points. Spend them.
Every speech-to-text trade-off is zero-sum — the most accurate model is rarely the cheapest or the fastest, and the most private one runs on your own hardware. So instead of rating everything “very important”, split 100 points across four priorities. Push one up and the others give way.
We rank the 26 cloud models and 15 on-device models HyperWhisper actually ships across macOS and Windows — not a general leaderboard. Pick your platform below and the list narrows to what you can run.
Your 100 points
40 / 20 / 30 / 10How often it gets a word wrong
How long you wait for the text
What a minute of audio costs you
Whether the audio leaves your machine
Best match for how you spent your points
Nemotron 3.5 Multilingual
On deviceNemotron · 1.3 GB download · runs entirely on your Mac or PC
- Word errors
- —
- Cost
- Free
- Per audio minute
- 1.2 s
- Match
- 82
no per-minute cost
Your audio never leaves the machine, so there is no per-minute cost and no region to worry about.
Every model, ranked
26 cloud · 13 on-device
| Model | WER | Cost | Per min | Why it scored | Match |
|---|---|---|---|---|---|
Nemotron 3.5 MultilingualOn deviceapp rating Nemotron · 1.3 GB | — | Free | 1.2 s | 82 | |
Nemotron 3.5 LatinOn deviceapp rating Nemotron · 350 MB | — | Free | 1.2 s | 82 | |
Universal-3.5 ProCloud AssemblyAI | 3% | 3.5 cr/min | 777 ms● | 78 | |
Gemini 3 FlashCloud Google Gemini | 2.9% | 3 cr/min | 1.1 s● | 78 | |
Universal-3 ProCloud AssemblyAI | 3.1% | 3.5 cr/min | 777 ms● | 77 | |
MAI-Transcribe 1.5Cloud Microsoft MAI-Transcribe 1.5 | 2.4% | 6 cr/min | 585 ms● | 77 | |
Apple SpeechOn deviceapp rating Apple · Built in | — | Free | 1.2 s | 74 | |
Whisper Large v3Cloud Groq Whisper | 4.1% | 1.85 cr/min | 201 ms● | 74 | |
Voxtral MiniCloud Mistral Voxtral | 3.8% | 3 cr/min | 521 ms● | 73 | |
Scribe v2Cloud ElevenLabs Scribe v2 | 2.2% | 9.83 cr/min | 1.1 s● | 72 | |
Whisper MediumOn deviceapp rating Whisper · 1.5 GB | — | Free | 1.5 s | 72 | |
Universal-2Cloud AssemblyAI | 3.8% | 2.5 cr/min | 777 ms● | 72 | |
Parakeet V3On devicesame weights Parakeet · 494 MB | 4.5% | Free | 1.2 s | 72 | |
GPT TranscribeCloud OpenAI Whisper | 3.3% | 4.5 cr/min | 997 ms● | 72 | |
Grok Speech-to-TextCloud Grok STT | 4% | 1.67 cr/min | 760 ms● | 72 | |
Whisper Large v3On devicesame weights Whisper · 3.1 GB | 4.1% | Free | 2.0 s | 71 | |
Whisper Large v2On devicesame weights Whisper · 2.9 GB | 4.1% | Free | 2.0 s | 71 | |
Whisper Large v3 TurboCloud Groq Whisper | 4.6% | 0.67 cr/min | 201 ms● | 71 | |
Gemini 2.5 ProCloud Google Gemini | 2.9% | 7.5 cr/min | 1.1 s● | 70 | |
Whisper Large v3 TurboOn devicesame weights Whisper · 809 MB | 4.6% | Free | 1.5 s | 69 | |
Gemini 3.1 ProCloud Google Gemini | 2.8% | 10 cr/min | 1.1 s● | 66 | |
Whisper SmallOn deviceapp rating Whisper · 466 MB | — | Free | 1.5 s | 64 | |
GPT-4o Mini TranscribeCloud OpenAI Whisper | 4.5% | 3 cr/min | 997 ms● | 63 | |
GPT-4o TranscribeCloud OpenAI Whisper | 4% | 6 cr/min | 997 ms● | 62 | |
WhisperCloud OpenAI Whisper | 4.1% | 6 cr/min | 997 ms● | 61 | |
Nova 2 GeneralCloudnot benchmarked Deepgram Nova 3 | — | 5.5 cr/min | 1.0 s● | 60 | |
Nova 2 MedicalCloudnot benchmarked Deepgram Nova 3 | — | 5.5 cr/min | 1.0 s● | 60 | |
Gemini 2.5 Flash LiteCloud Google Gemini | 5.2% | 0.8 cr/min | 1.1 s● | 60 | |
Gemini 2.5 FlashCloud Google Gemini | 5.1% | 2.4 cr/min | 1.1 s● | 58 | |
Whisper BaseOn deviceapp rating Whisper · 142 MB | — | Free | 1.2 s | 58 | |
Whisper TinyOn deviceapp rating Whisper · 39 MB | — | Free | 1.2 s | 58 | |
Async v5Cloud Soniox | 3.8% | 1.67 cr/min | 3.5 s● | 57 | |
Qwen3 ASROn deviceapp rating Qwen3 · 1.3 GB | — | Free | 1.5 s | 56 | |
Async v4Cloud Soniox | 3.9% | 1.67 cr/min | 3.5 s● | 56 | |
Parakeet V2On devicesame weights Parakeet · 474 MB | 6.4% | Free | 1.2 s | 54 | |
Nova 3 GeneralCloud Deepgram Nova 3 | 5.2% | 5.5 cr/min | 1.0 s● | 52 | |
Nova 3 MedicalCloud Deepgram Nova 3 | 5.2% | 5.5 cr/min | 1.0 s● | 52 | |
GPT Live TranscribeCloudnot benchmarked OpenAI Whisper | — | 17 cr/min | 997 ms● | 40 | |
Chirp 3Cloud Google Chirp 3 | 4.3% | 16 cr/min | 2.4 s● | 30 |
A ● marks a timing we measured ourselves from your region. Everything else is the published speed factor.
Cloud and on-device are different bargains
Every row on this page is one or the other, and the badge says which. The distinction is not a detail — it changes what you pay, what you wait for, and where your voice ends up.
Cloud models
Your audio is uploaded to HyperWhisper Cloud, which hands it to the provider and sends the text back. You get the strongest accuracy available and no download, and you pay per audio minute in credits. These are the models with published error rates, because a hosted model is something a third party can measure.
On-device models
The model is downloaded once and runs on your own Mac or PC. The audio never leaves the machine, there is nothing to pay per minute, and it works with no network at all. What you give up is the top of the accuracy table, and a few gigabytes of disk.
Where these numbers come from
- Accuracy
- Word error rate comes from the Artificial Analysis speech-to-text leaderboard — an independent third party, not us. We do not publish an accuracy claim of our own here. A model they have not measured shows a dash and is marked not benchmarked; it scores neutrally rather than being flattered by a number we invented.
- Why some on-device models borrow a cloud model's score
- Whisper and Parakeet are open weights. The leaderboard measured the same weight files we download, just running on someone else's hardware — so the error rate carries over and is marked same weights. Speed does not carry over, which is why those rows never borrow a speed figure. The remaining local models have no published benchmark at all and fall back to the app's own accuracy rating, marked app rating. Without that fallback the smallest, roughest models would score like frontier ones purely for being free and private.
- Cost
- Credits per audio minute, read from the same catalog the desktop apps read, so the price here is the price the app charges you. 1,000 credits are $1, which makes a model at 4.5 credits a minute $4.50 per 1,000 minutes of audio. On-device models cost nothing to run, so they score full marks on cost no matter how you weight it.
- Speed
- Where we have measured a provider from your nearest region, the number is our own median from the last 90 days — the same measurements behind the latency page, taken from short clips because that is what dictation is. Those rows carry a dot. Everything else falls back to the leaderboard's published speed factor. On-device timings are an estimate from the app's own speed rating and depend on your hardware, so read them as an ordering, not a promise.
- Privacy
- Three steps, by where the audio goes. On-device scores full marks: nothing is transmitted. A model you can use with your own API key scores half: the audio still reaches the vendor, but on your account rather than through us. A cloud-tier-only model scores lowest. Turn the privacy slider up and the ranking moves to on-device models, which is the honest answer to that question.
- Your region
- HyperWhisper Cloud runs in 17 regions and routes you to the nearest one. We ask a small endpoint which of those you are closest to so the speed numbers are the ones you would actually get; the answer is used once to pick a row and is never stored. You can override it with the dropdown. Off our edge — a local dev server, say — nothing is detected and the busiest region is selected instead.
- What this page will not tell you
- A ranking is not a recommendation for a job it has never seen. If you dictate medical or legal terms, transcribe heavy accents, or need speaker labels, the model that wins here on paper may still lose on your audio. Every model listed is switchable in Settings → Transcription, so the last word is a minute of your own speech, not a number on a website.