Which model should you use?
Tell us what matters and we will rank the 27 cloud and 19 on-device models HyperWhisper ships across macOS and Windows. Nothing here is a general leaderboard.
Your 100 points
40 / 20 / 30 / 10How often it gets a word wrong
How long you wait for the text
What a minute of audio costs you
Whether the audio leaves your machine
Best match for how you spent your points
Nemotron 3.5 Multilingual
On deviceNemotron · ~1.3 GB download · runs entirely on your Mac or PC
- Word errors
- —
- Cost
- Free
- Per audio minute
- 1.2 s
- Match
- 86
no per-minute cost
Your audio never leaves the machine, so there is no per-minute cost and no region to worry about.
Every model, ranked
27 cloud · 17 on-device
| Model | WER | Cost | Per audio min | Measured clip | Why it scored | Match |
|---|---|---|---|---|---|---|
Nemotron 3.5 MultilingualOn deviceapp rating Nemotron · ~1.3 GB · 30 languages | — | Free | 1.2 s | — | 86 | |
Nemotron 3.5 LatinOn deviceapp rating Nemotron · ~350 MB · 6 languages | — | Free | 1.2 s | — | 86 | |
Gemini 3.5 TranscribeCloud Google Gemini 3.5 Transcribe · language count not published | 2.6% | 5.5 cr/min | 977 ms | ●2.7 s | 80 | |
MAI-Transcribe 1.5Cloud Microsoft MAI-Transcribe · 42 languages | 2.4% | 6 cr/min | 577 ms | ●616 ms | 78 | |
Whisper Large v3On devicesame weights Whisper · 3.1 GB · 100 languages | 4.1% | Free | 2.0 s | — | 78 | |
Whisper Large v2On devicesame weights Whisper · 2.9 GB · 100 languages | 4.1% | Free | 2.0 s | — | 78 | |
Apple SpeechOn deviceapp rating Apple · Built-in · 23 languages | — | Free | 1.2 s | — | 78 | |
Whisper MediumOn deviceapp rating Whisper · 1.5 GB · 100 languages | — | Free | 1.5 s | — | 77 | |
Parakeet V3On devicesame weights Parakeet · 494 MB · 25 languages | 4.5% | Free | 1.2 s | — | 76 | |
Scribe v2Cloud ElevenLabs Scribe v2 · 99 languages | 2.2% | 9.83 cr/min | 1.3 s | ●851 ms | 75 | |
Grok Speech-to-TextCloud Grok STT · 25 languages | 4% | 1.67 cr/min | 511 ms | ●539 ms | 75 | |
Gemini 3 FlashCloud Google Gemini · language count not published | 2.9% | 3 cr/min | 4.0 s | ●929 ms | 74 | |
Universal-2Cloud AssemblyAI · 98 languages | 3.8% | 2.5 cr/min | 737 ms | ●728 ms | 74 | |
Whisper Large v3 TurboOn devicesame weights Whisper · 809 MB · 100 languages | 4.6% | Free | 1.5 s | — | 74 | |
GPT TranscribeCloud OpenAI Whisper · 100 languages | 3.3% | 4.5 cr/min | 1.8 s | ●1.4 s | 73 | |
Voxtral MiniCloud Mistral Voxtral · 13 languages | 3.8% | 3 cr/min | 1.0 s | ●498 ms | 73 | |
Whisper Large v3Cloud Groq Whisper · 100 languages | 4.1% | 1.85 cr/min | 901 ms | ●210 ms | 72 | |
Universal-3.5 ProCloud AssemblyAI · 98 languages | 3% | 3.5 cr/min | — | ●728 ms | 71 | |
Whisper Large v3 TurboCloud Groq Whisper · 100 languages | 4.6% | 0.67 cr/min | 741 ms | ●210 ms | 70 | |
Whisper Small (English)On deviceapp rating Whisper · 466 MB · 1 language | — | Free | 1.2 s | — | 70 | |
Whisper Base (English)On deviceapp rating Whisper · 142 MB · 1 language | — | Free | 1.2 s | — | 70 | |
Async v5Cloud Soniox · 60 languages | 3.8% | 1.67 cr/min | 3.4 s | ●3.6 s | 70 | |
Whisper Medium (English)On deviceapp rating Whisper · 1.5 GB · 1 language | — | Free | 1.5 s | — | 69 | |
Whisper SmallOn deviceapp rating Whisper · 466 MB · 100 languages | — | Free | 1.5 s | — | 69 | |
GPT-4o Mini TranscribeCloud OpenAI Whisper · 100 languages | 4.5% | 3 cr/min | 1.6 s | ●1.4 s | 65 | |
Gemini 2.5 ProCloud Google Gemini · language count not published | 2.9% | 7.5 cr/min | 4.8 s | ●929 ms | 65 | |
GPT-4o TranscribeCloud OpenAI Whisper · 100 languages | 4% | 6 cr/min | 1.8 s | ●877 ms | 64 | |
Gemini 2.5 Flash LiteCloud Google Gemini · language count not published | 5.2% | 0.8 cr/min | 1.1 s | ●929 ms | 63 | |
Whisper BaseOn deviceapp rating Whisper · 142 MB · 100 languages | — | Free | 1.2 s | — | 62 | |
Whisper TinyOn deviceapp rating Whisper · 39 MB · 100 languages | — | Free | 1.2 s | — | 62 | |
Whisper Tiny (English)On deviceapp rating Whisper · 39 MB · 1 language | — | Free | 1.2 s | — | 62 | |
WhisperCloud OpenAI Whisper · 100 languages | 4.1% | 6 cr/min | 2.3 s | ●1.4 s | 62 | |
Gemini 2.5 FlashCloud Google Gemini · language count not published | 5.1% | 2.4 cr/min | 1.1 s | ●929 ms | 61 | |
Qwen3 ASROn deviceapp rating Qwen3 · ~1.3 GB · 30 languages | — | Free | 1.5 s | — | 61 | |
Muse Voice Transcribe 1.0Cloudnot benchmarked Meta Muse Voice Transcribe · language count not published | — | 3 cr/min | — | ●2.6 s | 60 | |
MAI-Transcribe 2Cloudnot benchmarked Microsoft MAI-Transcribe · 60 languages | — | 1.67 cr/min | — | ●616 ms | 59 | |
Parakeet V2On devicesame weights Parakeet · 474 MB · 1 language | 6.4% | Free | 1.2 s | — | 58 | |
Nova 3 GeneralCloud Deepgram Nova 3 · 64 languages | 5.2% | 5.5 cr/min | 361 ms | ●1.1 s | 57 | |
Nova 3 MedicalCloud Deepgram Nova 3 · 64 languages | 5.2% | 5.5 cr/min | 361 ms | ●1.1 s | 57 | |
Nova 2 GeneralCloudnot benchmarked Deepgram Nova 3 · 64 languages | — | 5.5 cr/min | — | ●1.1 s | 55 | |
Nova 2 MedicalCloudnot benchmarked Deepgram Nova 3 · 64 languages | — | 5.5 cr/min | — | ●1.1 s | 55 | |
Gemini 3.1 ProCloud Google Gemini · language count not published | 2.8% | 10 cr/min | 8.8 s | ●929 ms | 52 | |
Gemini 3.5 Transcribe LiveCloudnot benchmarked Google Gemini 3.5 Transcribe · language count not published | — | 9.6 cr/min | — | ●2.7 s | 48 | |
GPT Live TranscribeCloudnot benchmarked OpenAI Whisper · 100 languages | — | 17 cr/min | — | ●1.4 s | 35 |
The two speed columns are different questions. Per audio min is how long a minute of audio takes, from the leaderboard's published speed factor — or, for on-device rows, estimated from the app's own speed rating. Measured clip, marked ●, is our own median wall time for one dictation clip under 10 seconds from your region. A short clip is mostly round trip rather than decoding, so the two do not convert into one another — and the ranking uses only the first, so a model we have measured is never compared against one we have not on a different footing.
Cloud and on-device are different bargains
Every row on this page is one or the other, and the badge says which. The distinction is not a detail — it changes what you pay, what you wait for, and where your voice ends up.
Cloud models
Your audio is uploaded to HyperWhisper Cloud, which hands it to the provider and sends the text back. You get the strongest accuracy available and no download, and you pay per audio minute in credits. These are the models with published error rates, because a hosted model is something a third party can measure.
On-device models
The model is downloaded once and runs on your own Mac or PC. The audio never leaves the machine, there is nothing to pay per minute, and it works with no network at all. What you give up is the top of the accuracy table, and a few gigabytes of disk.
Where these numbers come from
- Accuracy
- Word error rate comes from the Artificial Analysis speech-to-text leaderboard — an independent third party, not us. We do not publish an accuracy claim of our own here. A model they have not measured shows a dash and is marked not benchmarked; it scores neutrally rather than being flattered by a number we invented.
- Why some on-device models borrow a cloud model's score
- Whisper and Parakeet are open weights. The leaderboard measured the same weight files we download, just running on someone else's hardware — so the error rate carries over and is marked same weights. Speed does not carry over, which is why those rows never borrow a speed figure. The remaining local models have no published benchmark at all and fall back to the app's own accuracy rating, marked app rating. Without that fallback the smallest, roughest models would score like frontier ones purely for being free and private.
- Cost
- Credits per audio minute, read from the same catalog the desktop apps read, so the price here is the price the app charges you. 1,000 credits are $1, which makes a model at 4.5 credits a minute $4.50 per 1,000 minutes of audio. On-device models cost nothing to run, so they score full marks on cost no matter how you weight it.
- Speed
- Two numbers, and the difference matters. Per audio minute is how long a minute of speech takes to come back, from the leaderboard's published speed factor; on-device rows estimate it from the app's own speed rating, so read those as an ordering, not a promise. That is the number the ranking uses, because it is the only one we hold for every model. Measured clip is our own median from the last 90 days — the same measurements behind the latency page, taken from clips under ten seconds because that is what dictation is. It is a round trip for one clip, not a throughput: most of it is the network and the provider's fixed overhead rather than decoding, and we do not store the clip's length, so it cannot be scaled up to a minute. Feeding it into the ranking anyway would punish the providers we happen to know most about, so it sits in its own column instead.
- Privacy
- Three steps, by where the audio goes. On-device scores full marks: nothing is transmitted. A model you can use with your own API key scores half: the audio still reaches the vendor, but on your account rather than through us. A cloud-tier-only model scores lowest. Turn the privacy slider up and the ranking moves to on-device models, which is the honest answer to that question.
- Your region
- HyperWhisper Cloud runs in 17 regions and routes you to the nearest one. We ask a small endpoint which of those you are closest to so the speed numbers are the ones you would actually get; the answer is used once to pick a row and is never stored. You can override it with the dropdown. Off our edge — a local dev server, say — nothing is detected and the busiest region is selected instead.
- What this page will not tell you
- A ranking is not a recommendation for a job it has never seen. If you dictate medical or legal terms, transcribe heavy accents, or need speaker labels, the model that wins here on paper may still lose on your audio. Every model listed is switchable in Settings → Transcription, so the last word is a minute of your own speech, not a number on a website.