Skip to main content
Streaming transcription converts your speech to text in real time, word by word. It is separate from the mode-based recording system. You start it with its own keyboard shortcut. It uses its own provider and language settings.

Turn On Streaming

Streaming is off by default. The streaming shortcut and the other streaming settings do not operate until you turn on streaming.
Open the Streaming section in the sidebar. Then turn on the Enable Streaming toggle.The shortcut field and the other options then appear below the toggle.

Keyboard Shortcut

After you turn on streaming, set the shortcut that starts it.
The default shortcut is ⌥ ⇧ Space (Option + Shift + Space). To change it, click the shortcut field in the Streaming sidebar section. Then press the new key combination. Press the shortcut to start streaming. Press it again to stop streaming.If the shortcut conflicts with another HyperWhisper shortcut, a warning appears. HyperWhisper does not save the new shortcut until you remove the conflict.

Select a Provider

The Engine section is visible only after you turn on streaming. Open this section, then select a provider. HyperWhisper Cloud is the default provider on macOS and Windows. It needs no provider API key — only your HyperWhisper Cloud account key and a credit balance. For each other cloud provider, you must add an API key in Model Library → API Keys. The provider does not operate without this key. If a key is missing or not valid, a warning appears below the provider picker. Gemini 3.5 Transcribe uses its own key slot, and not the Google Gemini slot. The two are different Google products. Put your Google key in the Gemini 3.5 Transcribe slot in Model Library → API Keys. See API Keys. Parakeet and Nemotron 3.5 run fully on your device. After the model download is complete, they need no network connection. The provider list is not the same on each platform: On Linux, the two on-device providers run through the same Parakeet engine process that Linux uses for batch transcription.
On Linux, the provider picker shows the internal identifiers, not the display names: deepgram, elevenlabs, openai, grok (xAI), geminiTranscribe (Gemini 3.5 Transcribe), hyperwhisper (HyperWhisper Cloud), parakeetLocal, and nemotronLocal. The default is deepgram, and not HyperWhisper Cloud. Linux reads the provider key from the Credentials page and never keeps it in the settings file.You can also select a local model in the Unified model library. Select an installed streaming model, then click Use for live transcription. HyperWhisper turns on streaming and sets the provider and the model for you.
HyperWhisper sends the entries from the Vocabulary section to HyperWhisper Cloud, Deepgram, xAI and Gemini 3.5 Transcribe. For Deepgram you must set an explicit language; xAI and Gemini 3.5 Transcribe accept the terms with any language, including Automatic. For HyperWhisper Cloud the rule follows the live engine: Deepgram Nova 3 needs an explicit language, and Gemini 3.5 Transcribe does not. ElevenLabs and OpenAI do not support vocabulary boosting in their streaming API. If you have vocabulary entries and select one of those two providers, a warning appears. Text replacements apply during streaming only on macOS with an on-device provider (Parakeet or Nemotron 3.5). Cloud streaming providers receive only the vocabulary entries that have no replacement, as boosting hints. This is the same rule on macOS and Windows. Linux sends every vocabulary word as a boosting hint, and it does not remove a word that has a replacement.

HyperWhisper Cloud Live Engine

HyperWhisper Cloud sends your live audio to one of two engines. A second picker appears below the provider picker when the provider is HyperWhisper Cloud. It is hidden for every other provider. The picker changes the engine only. It does not change your account key, and it does not change the Language setting. The engine that you select sets the credit rate. For the rates of the batch engines, see Cloud Credits.
Open the Streaming section in the sidebar. The engine picker is below the provider picker.

Gemini 3.5 Transcribe Options

Gemini 3.5 Transcribe has one live model, thus there is no model picker. HyperWhisper sends your vocabulary entries with any language, and also with the language set to Automatic. Google keeps the socket open after your last word. HyperWhisper waits 5 seconds for the last part of the text, then it closes the session.

Deepgram Options

When you select Deepgram, two more settings appear.

Model

Select Nova 3 General (the default) or Nova 3 Medical. The medical model is tuned for healthcare terms and clinical language. It transcribes English only. Set the streaming language to English when you select it.

Fast Formatting

This setting is on by default on macOS and Windows. When it is on, Deepgram sends smart-formatted results immediately and does not wait for more context. Thus the delay before the words appear on screen is smaller. When it is off, Deepgram waits for more context before it sets the punctuation and the number format. This gives slightly more accurate formatting, but it adds latency.
On Linux, the setting is the Enable provider fast formatting checkbox, and it is off by default.

On-Device Providers

Parakeet

Parakeet is the open-weight speech model from NVIDIA. It runs on your device with the integrated on-device engine.
  • Parakeet V3 (Multilingual) — it supports 25 European languages. This is the default.
  • Parakeet V2 (English) — English only. It has the highest recall.
You must download the model before the first use. When the model is not on disk, an Install button and a download progress indicator appear in the Engine section. After the installation, click Manage to open the Model Library.

Nemotron 3.5

Nemotron 3.5 is the streaming ASR model family from NVIDIA. It also runs fully on your device.
  • Nemotron 3.5 (Multilingual) — the model picker shows approximately 30 languages. These include Chinese, Japanese, Korean, and Arabic. This is the default variant.
  • Nemotron 3.5 (Latin) — approximately 6 Latin-script languages. It is faster.
The same installation steps apply. You must download the variant that you select before streaming can start.

On Linux

Linux has no Engine card and no Install button in the live-transcription section. Select parakeetLocal or nemotronLocal in the provider picker, then set the model in the Provider model (optional) box. Download the model first, in the Unified model library.
The Nemotron Latin variant is not available for live transcription on Linux.

Language

The Language setting tells the provider which language to expect. An explicit language usually gives better accuracy. For Deepgram, an explicit language is necessary for vocabulary boosting. The default language on macOS and Windows is English. On Linux the default is Automatic.
The language picker appears in the Language section, below the Engine card. The available choices depend on the provider that you select. For cloud providers, the Automatic option lets the server detect the language from your audio. When you select Automatic, vocabulary boosting is off for HyperWhisper Cloud and Deepgram. It stays on for xAI.For Parakeet V2, the language is always English, and the picker is off. For Parakeet V3 and the two Nemotron variants, the picker shows only the languages that the model supports.

If the Connection Drops

A streaming session uses one WebSocket connection to the provider. If that connection stops without warning, HyperWhisper keeps the microphone open and connects again automatically.
  • The status indicator changes to Reconnecting. On macOS the indicator is yellow and pulses. On Windows it is orange and pulses. Read Recording Window for the full list of the indicator colors.
  • HyperWhisper makes a maximum of 3 attempts. On macOS, it waits 1 second before each attempt. On Windows, it waits 500 ms before the first attempt, and more before each subsequent attempt. On Linux, it waits 250 ms before the first attempt and 500 ms before the second.
  • If all 3 attempts fail, HyperWhisper stops the session and shows an error. Start the session again with your streaming shortcut.
  • On macOS, a connection that stayed up for 60 seconds or more before it dropped sets the count of the attempts back to zero. A long session thus does not stop after 3 separate network problems. Drops that occur one after the other in a few seconds continue to add to the count.
Some faults cannot improve with another attempt. If the provider reports an exhausted credit balance, a key that is not valid, or a permission error, HyperWhisper stops immediately and shows that message. It makes no attempt to connect again, because the same fault occurs each time. Correct the condition, then start the session again.
To start a streaming session with HyperWhisper Cloud, you need a credit balance of approximately 30 seconds of audio. The server calculates this balance with the rate of the live engine that you selected, thus Gemini 3.5 Transcribe needs more credits than Deepgram Nova 3. Below that balance, the server refuses the session before it starts, and the app tells you that credits are exhausted. Read HyperWhisper Cloud Credits to add credits.

How Streaming Differs from Modes

Standard recording in HyperWhisper uses your transcription modes. Each mode can have its own provider, language, and post-processing settings. Streaming is a separate path that is always available. Its settings apply to every streaming session. Streaming does not use the mode settings, and it does not apply AI post-processing. The Remove Filler Words Text Output setting also applies to streaming. It removes the filler words from each confirmed chunk when the chunk arrives. On Windows, streaming removes filler words only when the streaming language is English. For batch transcription, Windows removes filler words in all languages. Use streaming when you want the words to appear on screen in real time. Use a mode when you want a complete transcript after you speak. A mode can also apply post-processing to that transcript.