HyperWhisper is now fully open source · Now open source · Learn more

  • HyperWhisper Logo

    HyperWhisper

    • Features
    • Cloud
    • Choose a model
    • Latency
    • FAQ

Which model should you use?

Tell us what matters and we will rank the 27 cloud and 19 on-device models HyperWhisper ships across macOS and Windows. Nothing here is a general leaderboard.

Your 100 points

40 / 20 / 30 / 10
40

How often it gets a word wrong

20

How long you wait for the text

30

What a minute of audio costs you

10

Whether the audio leaves your machine

Your platform
Languages you dictate
Must have
Closest region

Best match for how you spent your points

Nemotron 3.5 Multilingual

On device

Nemotron · ~1.3 GB download · runs entirely on your Mac or PC

Word errors
—
Cost
Free

no per-minute cost

Per audio minute
1.2 s
Match
86

Your audio never leaves the machine, so there is no per-minute cost and no region to worry about.

Every model, ranked

27 cloud · 17 on-device

ModelWERCostPer audio minMeasured clipWhy it scoredMatch
Nemotron 3.5 MultilingualOn deviceapp rating
Nemotron · ~1.3 GB · 30 languages
—Free1.2 s—
86
Nemotron 3.5 LatinOn deviceapp rating
Nemotron · ~350 MB · 6 languages
—Free1.2 s—
86
Gemini 3.5 TranscribeCloud
Google Gemini 3.5 Transcribe · language count not published
2.6%5.5 cr/min977 ms●2.7 s
80
MAI-Transcribe 1.5Cloud
Microsoft MAI-Transcribe · 42 languages
2.4%6 cr/min577 ms●616 ms
78
Whisper Large v3On devicesame weights
Whisper · 3.1 GB · 100 languages
4.1%Free2.0 s—
78
Whisper Large v2On devicesame weights
Whisper · 2.9 GB · 100 languages
4.1%Free2.0 s—
78
Apple SpeechOn deviceapp rating
Apple · Built-in · 23 languages
—Free1.2 s—
78
Whisper MediumOn deviceapp rating
Whisper · 1.5 GB · 100 languages
—Free1.5 s—
77
Parakeet V3On devicesame weights
Parakeet · 494 MB · 25 languages
4.5%Free1.2 s—
76
Scribe v2Cloud
ElevenLabs Scribe v2 · 99 languages
2.2%9.83 cr/min1.3 s●851 ms
75
Grok Speech-to-TextCloud
Grok STT · 25 languages
4%1.67 cr/min511 ms●539 ms
75
Gemini 3 FlashCloud
Google Gemini · language count not published
2.9%3 cr/min4.0 s●929 ms
74
Universal-2Cloud
AssemblyAI · 98 languages
3.8%2.5 cr/min737 ms●728 ms
74
Whisper Large v3 TurboOn devicesame weights
Whisper · 809 MB · 100 languages
4.6%Free1.5 s—
74
GPT TranscribeCloud
OpenAI Whisper · 100 languages
3.3%4.5 cr/min1.8 s●1.4 s
73
Voxtral MiniCloud
Mistral Voxtral · 13 languages
3.8%3 cr/min1.0 s●498 ms
73
Whisper Large v3Cloud
Groq Whisper · 100 languages
4.1%1.85 cr/min901 ms●210 ms
72
Universal-3.5 ProCloud
AssemblyAI · 98 languages
3%3.5 cr/min—●728 ms
71
Whisper Large v3 TurboCloud
Groq Whisper · 100 languages
4.6%0.67 cr/min741 ms●210 ms
70
Whisper Small (English)On deviceapp rating
Whisper · 466 MB · 1 language
—Free1.2 s—
70
Whisper Base (English)On deviceapp rating
Whisper · 142 MB · 1 language
—Free1.2 s—
70
Async v5Cloud
Soniox · 60 languages
3.8%1.67 cr/min3.4 s●3.6 s
70
Whisper Medium (English)On deviceapp rating
Whisper · 1.5 GB · 1 language
—Free1.5 s—
69
Whisper SmallOn deviceapp rating
Whisper · 466 MB · 100 languages
—Free1.5 s—
69
GPT-4o Mini TranscribeCloud
OpenAI Whisper · 100 languages
4.5%3 cr/min1.6 s●1.4 s
65
Gemini 2.5 ProCloud
Google Gemini · language count not published
2.9%7.5 cr/min4.8 s●929 ms
65
GPT-4o TranscribeCloud
OpenAI Whisper · 100 languages
4%6 cr/min1.8 s●877 ms
64
Gemini 2.5 Flash LiteCloud
Google Gemini · language count not published
5.2%0.8 cr/min1.1 s●929 ms
63
Whisper BaseOn deviceapp rating
Whisper · 142 MB · 100 languages
—Free1.2 s—
62
Whisper TinyOn deviceapp rating
Whisper · 39 MB · 100 languages
—Free1.2 s—
62
Whisper Tiny (English)On deviceapp rating
Whisper · 39 MB · 1 language
—Free1.2 s—
62
WhisperCloud
OpenAI Whisper · 100 languages
4.1%6 cr/min2.3 s●1.4 s
62
Gemini 2.5 FlashCloud
Google Gemini · language count not published
5.1%2.4 cr/min1.1 s●929 ms
61
Qwen3 ASROn deviceapp rating
Qwen3 · ~1.3 GB · 30 languages
—Free1.5 s—
61
Muse Voice Transcribe 1.0Cloudnot benchmarked
Meta Muse Voice Transcribe · language count not published
—3 cr/min—●2.6 s
60
MAI-Transcribe 2Cloudnot benchmarked
Microsoft MAI-Transcribe · 60 languages
—1.67 cr/min—●616 ms
59
Parakeet V2On devicesame weights
Parakeet · 474 MB · 1 language
6.4%Free1.2 s—
58
Nova 3 GeneralCloud
Deepgram Nova 3 · 64 languages
5.2%5.5 cr/min361 ms●1.1 s
57
Nova 3 MedicalCloud
Deepgram Nova 3 · 64 languages
5.2%5.5 cr/min361 ms●1.1 s
57
Nova 2 GeneralCloudnot benchmarked
Deepgram Nova 3 · 64 languages
—5.5 cr/min—●1.1 s
55
Nova 2 MedicalCloudnot benchmarked
Deepgram Nova 3 · 64 languages
—5.5 cr/min—●1.1 s
55
Gemini 3.1 ProCloud
Google Gemini · language count not published
2.8%10 cr/min8.8 s●929 ms
52
Gemini 3.5 Transcribe LiveCloudnot benchmarked
Google Gemini 3.5 Transcribe · language count not published
—9.6 cr/min—●2.7 s
48
GPT Live TranscribeCloudnot benchmarked
OpenAI Whisper · 100 languages
—17 cr/min—●1.4 s
35

The two speed columns are different questions. Per audio min is how long a minute of audio takes, from the leaderboard's published speed factor — or, for on-device rows, estimated from the app's own speed rating. Measured clip, marked ●, is our own median wall time for one dictation clip under 10 seconds from your region. A short clip is mostly round trip rather than decoding, so the two do not convert into one another — and the ranking uses only the first, so a model we have measured is never compared against one we have not on a different footing.

Cloud and on-device are different bargains

Every row on this page is one or the other, and the badge says which. The distinction is not a detail — it changes what you pay, what you wait for, and where your voice ends up.

Cloud models

Your audio is uploaded to HyperWhisper Cloud, which hands it to the provider and sends the text back. You get the strongest accuracy available and no download, and you pay per audio minute in credits. These are the models with published error rates, because a hosted model is something a third party can measure.

On-device models

The model is downloaded once and runs on your own Mac or PC. The audio never leaves the machine, there is nothing to pay per minute, and it works with no network at all. What you give up is the top of the accuracy table, and a few gigabytes of disk.

Where these numbers come from

Accuracy
Word error rate comes from the Artificial Analysis speech-to-text leaderboard — an independent third party, not us. We do not publish an accuracy claim of our own here. A model they have not measured shows a dash and is marked not benchmarked; it scores neutrally rather than being flattered by a number we invented.
Why some on-device models borrow a cloud model's score
Whisper and Parakeet are open weights. The leaderboard measured the same weight files we download, just running on someone else's hardware — so the error rate carries over and is marked same weights. Speed does not carry over, which is why those rows never borrow a speed figure. The remaining local models have no published benchmark at all and fall back to the app's own accuracy rating, marked app rating. Without that fallback the smallest, roughest models would score like frontier ones purely for being free and private.
Cost
Credits per audio minute, read from the same catalog the desktop apps read, so the price here is the price the app charges you. 1,000 credits are $1, which makes a model at 4.5 credits a minute $4.50 per 1,000 minutes of audio. On-device models cost nothing to run, so they score full marks on cost no matter how you weight it.
Speed
Two numbers, and the difference matters. Per audio minute is how long a minute of speech takes to come back, from the leaderboard's published speed factor; on-device rows estimate it from the app's own speed rating, so read those as an ordering, not a promise. That is the number the ranking uses, because it is the only one we hold for every model. Measured clip is our own median from the last 90 days — the same measurements behind the latency page, taken from clips under ten seconds because that is what dictation is. It is a round trip for one clip, not a throughput: most of it is the network and the provider's fixed overhead rather than decoding, and we do not store the clip's length, so it cannot be scaled up to a minute. Feeding it into the ranking anyway would punish the providers we happen to know most about, so it sits in its own column instead.
Privacy
Three steps, by where the audio goes. On-device scores full marks: nothing is transmitted. A model you can use with your own API key scores half: the audio still reaches the vendor, but on your account rather than through us. A cloud-tier-only model scores lowest. Turn the privacy slider up and the ranking moves to on-device models, which is the honest answer to that question.
Your region
HyperWhisper Cloud runs in 17 regions and routes you to the nearest one. We ask a small endpoint which of those you are closest to so the speed numbers are the ones you would actually get; the answer is used once to pick a row and is never stored. You can override it with the dropdown. Off our edge — a local dev server, say — nothing is detected and the busiest region is selected instead.
What this page will not tell you
A ranking is not a recommendation for a job it has never seen. If you dictate medical or legal terms, transcribe heavy accents, or need speaker labels, the model that wins here on paper may still lose on your audio. Every model listed is switchable in Settings → Transcription, so the last word is a minute of your own speech, not a number on a website.
HyperWhisper LogoHyperWhisper

Write 5x faster with AI-powered voice transcription for macOS, Windows & Linux.

Product

  • Features
  • Pricing
  • Roadmap

Resources

  • Help Center
  • Customer Portal
  • Older Versions
  • Blog
  • Open Source

Company

  • About
  • Support

Legal

  • Privacy Policy
  • Terms of Service
  • Refund Policy
  • Data Privacy

© 2026 HyperWhisper. All rights reserved.