Skip to main content
HyperWhisper stores a transcription language for each mode. The engine can also detect the language from the audio. This detection works best with long recordings.

Setting the Language for a Mode

HyperWhisper stores the language per mode. Each mode in your list can use a different language. Keep one mode for each language that you use frequently. Then you change no setting between tasks.
1

Open the mode editor

Click the HyperWhisper menu bar icon. The main window opens. Click Modes in the sidebar. Click the mode that you want to edit. To make a new mode, click Create Mode.
2

Change the language

In the mode editor, find the Language row in the Transcription section. Click the picker to see all languages. Popular languages are at the top. The other languages are below them, in alphabetical order.
3

Save

Click Save. HyperWhisper stores the language with the mode. The mode uses this language for the next recording.
The default selection is Automatic. With this selection, the engine detects the language from the audio. Automatic Language Detection below gives the correct uses and the limits of this function.
One separate mode for each frequent language is the best method. The mode picker and the keyboard shortcut change modes immediately. Each mode keeps its own language, so you change no settings between sessions.

Automatic Language Detection

If you select Automatic, HyperWhisper sends no language hint to the transcription engine. The engine finds the language from the audio content. Auto-detect works best when:
  • The recording has 10–15 seconds or more of clear speech
  • You speak one language in the full recording
  • The model has much training data for the language that you speak
Auto-detect is unreliable when:
  • The recording is short (less than 10–15 seconds). The engine can have too little speech to find the language with confidence
  • You change language during the recording
Short recordings with auto-detect can give incorrect text, empty results, or the wrong language. For most tasks, a specific language gives better and more constant results. Best Practices gives more information.
Some models and cloud providers do not support auto-detect. If a model supports only English, the mode editor hides the language picker. The engine then transcribes only in English.

English-Only Models

Some local models support only English. For these models, the mode editor hides the language picker and gives no language detection:
  • Whisper .en variants (for example, base.en, small.en) — English-only builds of Whisper with better accuracy for English
  • Parakeet V2 — the English-only Parakeet model from NVIDIA with the highest recall
If you select one of these models, the mode editor shows a notice in place of the picker. The mode then transcribes only in English.

Supported Languages by Model

The number of available languages depends on the model or the provider that you select. For cloud providers, the language picker shows only the languages that the selected model supports. If you change to a model with fewer languages, and that model does not support your language, the mode changes to the first available language. The Model Library gives the full list for each model. Popular languages are at the top of the picker on both platforms: English, Japanese, Spanish, Chinese, Chinese (Traditional), Dutch, Hindi, Russian, Korean, Italian, Ukrainian, Polish, Portuguese, Greek, Czech, Swedish, Norwegian, Danish, and Indonesian. The other languages come after them, in alphabetical order.

How Language Affects Other Features

The language of a mode affects more than the transcription engine.

Vocabulary Boosting

HyperWhisper sends the words from Vocabulary to cloud providers as hints. These hints increase the accuracy for proper nouns, technical terms, and the words of your field. Deepgram Nova-3: if the language is Automatic, Deepgram ignores the keyterm parameter and the vocabulary hints. For vocabulary boosting with Deepgram Nova-3, select a specific language in the mode. If you select Nova-3 with auto-detect, the mode editor shows a notice. Some providers do not support vocabulary hints. For these providers, the mode editor shows a notice.

AI Post-Processing

If you enable AI post-processing, a language model improves the transcription result. The English Spelling setting (American, British, Australian, or Canadian) is available only when the mode language is English. For other languages, the spelling follows the language of the mode.

Engine Availability

If a model does not support your language, the picker changes to the first language in the list of that model. HyperWhisper makes this change when you select the model. The Deepgram medical models are an exception. The picker filters the languages by accuracy tier, not by model, so all the languages of the Deepgram tier stay in the list. But Nova 3 Medical and Nova 2 Medical transcribe English only. If you select a medical model with a different language, HyperWhisper Cloud removes the language and Deepgram detects it from the audio.

Language Codes for Cloud Providers

Each provider has its own spelling of a language code. For a request through HyperWhisper Cloud, the server changes the code of the mode into the code that the provider accepts. Examples:
  • Google Gemini 3.5 Transcribe accepts a bare code such as en and a full locale such as en-US. HyperWhisper Cloud sends the code of the mode without a change.
  • Azure MAI-Transcribe lists Norwegian as nb. HyperWhisper Cloud sends nb for Norwegian.
  • ElevenLabs Scribe lists Tagalog as fil, and Mandarin as cmn.
If the provider has no code for your language, HyperWhisper Cloud sends no language. The provider then detects the language from the audio. A transcription with automatic detection is better than a failed request.
This repair is part of HyperWhisper Cloud. It applies to the providers that HyperWhisper Cloud operates, not to a request that you send with your own API key.

Streaming Language

Streaming transcription has its own language setting. This setting is not the language of a mode.
The Streaming section of the sidebar contains the streaming language. The default is English. You can change it to a supported language or to Automatic. The picker shows only the languages of the streaming provider and model that you selected. For example, Parakeet V3 streaming limits the picker to its 25 languages.
The auto-detect limits also apply to streaming. If the streaming language is Automatic, short inputs and single sentences can give inconsistent results.
The Deepgram vocabulary boosting limit also applies to streaming. If the streaming language is Automatic, Deepgram Nova-3 receives no vocabulary hints.