Skip to main content
Three choices control performance in HyperWhisper. The first is where the transcription runs: in the cloud or on your device. The second is which model you select for your hardware. The third is a group of settings for startup time, audio processing, and storage. This page explains each one.

Select your transcription strategy

Cloud — fastest results, highest accuracy

HyperWhisper Cloud routes your audio to the best available providers. It needs no setup. The server returns results faster than most on-device models can load and process the same audio. HyperWhisper Cloud also includes Azure MAI-Transcribe and Google Gemini 3.5 Transcribe. Both have a High accuracy rating. Several other providers are available. For the full list, see Choosing a Provider. Cost note: HyperWhisper Cloud detects silence and empty audio automatically. If a recording contains no speech, it costs 0 credits. Pauses, dead air, and empty recordings that you start by accident cost nothing. In a normal push-to-talk workday, you pay only for the minutes that you spoke.

On-device — offline, private, no per-minute cost

Local models run only on your machine. Your audio never leaves your device. After you download the model, there is no per-minute charge. Local models are a little less accurate than the top cloud tiers. Each model needs one download of 350 MB to 3.1 GB.
On-device streaming transcription (Parakeet V2/V3, Nemotron 3.5 Streaming) works on macOS only. On Windows, streaming supports cloud providers only: HyperWhisper Cloud, Deepgram, ElevenLabs, OpenAI, and Grok. Windows has no local-model streaming. For the full platform matrix, see the Models page.

Select the correct model size for your hardware

Apple Silicon Macs (M1 and later)

Local models use Metal GPU and Neural Engine acceleration. The Small Whisper model (466 MB on macOS, ~2 GB VRAM) runs in realtime, also on an M1 Air. It is the recommended starting point for most users. For English-only work, Parakeet V2 (474 MB) is usually faster than Whisper models of the same size. Nemotron 3.5 Multilingual (~1.3 GB) is the only local model that supports languages outside Europe, such as Chinese, Japanese, Korean, and Arabic.

Intel Macs

Local models work on Intel Macs, but they use the CPU only. There is no GPU or Neural Engine acceleration. Start with Whisper Tiny (~39 MB) or Whisper Base (148 MB). If these models are too slow or not accurate enough, change to HyperWhisper Cloud. The server then does all the computation.

Windows x64

Local transcription on Windows uses Vulkan (Whisper) and DirectML (Parakeet) for GPU acceleration. Vulkan works with all vendors. Any recent NVIDIA, AMD, or Intel GPU with Vulkan drivers is sufficient. If the app finds no compatible GPU, the model changes to the CPU automatically. The model still works, but it is slower.
On ARM64 and Snapdragon Windows devices, Whisper requires Windows 11 or Windows Server 2022 or newer. Whisper is not available on ARM64 devices that still run Windows 10. On these devices, the local options are Parakeet V2 (English) and Parakeet V3 (25 European languages). Both use DirectML acceleration.

Whisper size ladder

Small models are faster, but they are less accurate. Large models are slower, but they are better with difficult audio. If your GPU does not have sufficient VRAM, the model changes to the CPU automatically.
The recommended VRAM figures for Whisper models come from the Windows build. macOS uses the unified memory of Apple Silicon. Unified memory has no separate VRAM.
If you always dictate in one language, the English-only Whisper models (.en) give slightly better results at the same model size. In exchange, you lose multilingual support.

Reduce push-to-talk startup latency

Enable “Keep Microphone Warm”

A cold microphone start is the largest source of push-to-talk delay. Bluetooth headsets show this delay most, because they can need one second or more to change into call mode. Keep Microphone Warm holds a quiet idle audio session open between recordings. The microphone is then ready when you press your shortcut.
Enable it in Settings → Sound → Keep Microphone Warm.
When this setting is on, macOS always shows the orange microphone indicator in the menu bar, also when you do not record. Bluetooth headsets can also stay in their lower-quality call audio profile instead of a change back to stereo. If one of these effects is a problem for you, you can turn the setting off.

Deepgram Fast Formatting (streaming)

When you use Deepgram for streaming transcription, Fast Formatting is on by default. It returns smart-formatted results immediately, and does not wait for the nearby context. Words then appear on screen sooner. If you turn it off, punctuation and number formatting become slightly more accurate, but latency increases. Unless formatting precision is more important to you than speed, leave it on.

Optimize file transcription

Enable VAD (Voice Activity Detection)

Remove silence before transcription (in Settings → Sound) examines the clip after you stop the recording. It uses an AI voice-detection model (Silero VAD) to remove the silence at the start and at the end. It then sends the audio to the provider.Why this setting helps:
  • It sends less audio to cloud providers, which can lower API costs.
  • It makes transcription faster, above all for short clips with long pauses at the start or the end.
  • It can improve accuracy, because it removes segments that contain only noise.
If your recordings always start and end with speech, leave it off. VAD then adds a small processing step with no benefit.

Sample rate

The default audio sample rate for file transcription is 16 000 Hz. This rate is the standard input rate for Whisper models, and most cloud providers expect it. A higher rate gives no benefit for transcription.
This setting applies to file and push-to-talk recordings on macOS. Streaming transcription uses a fixed rate from the streaming provider, not this setting.

Post-processing performance

Post-processing is an optional second step after transcription. It removes filler words, corrects punctuation, and applies formatting. Post-processing also adds latency, which is important when speed is your priority.

Local Gemma models (offline, no extra cost)

Local Gemma post-processing uses Metal GPU acceleration on Apple Silicon Macs (M1 and later). Intel Macs do not support local LLM post-processing. On an Intel Mac, use a cloud post-processing provider.Start with Gemma 4 E2B. It fits in 4 GB and does most cleanup tasks well.

Cloud post-processing

If local Gemma is too large for your machine, use a cloud post-processing provider. These providers include HyperWhisper Cloud, OpenAI, Claude, Gemini, Groq, and more. They need no local storage. The Model Library shows a speed rating and an accuracy rating for each one.

Storage and disk space

Keep audio files (macOS) is off by default. When transcription succeeds, the app deletes the audio to save disk space. If you want playback from History, or you want to do failed transcriptions again, turn this setting on. Windows has no equivalent setting. Windows deletes audio automatically after a number of days that you set. For details, see History. Max recording duration on macOS is 3600 seconds (1 hour) by default. You can change it with the presets (20 min / 1 hr / 2 hr / Off) in Settings → Storage. Windows has a fixed hard stop at 20 minutes (1200 s). You cannot change this limit.

Quick selection guide

Find your goal in the table below.
  • Models — full model library, VRAM requirements, and speed/accuracy ratings
  • System Requirements — hardware specifications for each platform
  • Providers — cloud tier pricing, silence-free billing, and cost examples
  • Best Practices — accuracy tips for vocabulary, microphone hardware, and environment