Use these guidelines in all your recordings. They make your transcriptions more accurate and more reliable.
Specify Your Language
Do not use “Auto-detect” for short recordings. The transcription provider needs sufficient audio to identify the language that you speak. A recording of less than 10–15 seconds can give too little speech data for reliable detection.
Auto-detect on a short recording causes these problems:
- Nonsense output – The model can transcribe your speech as a different language and give unreadable text
- Empty results – If the provider cannot identify the language, it returns no text
- Inconsistent behavior – The same short phrase can transcribe correctly one time and fail the next time
Create a separate mode for each language that you use regularly. This gives you the best accuracy. You can still change language quickly with the mode picker or a keyboard shortcut.
Speak Clearly and Naturally
- Keep the same distance from your microphone
- Remove background noise when you can
- Speak at a natural speed, because fast speech decreases accuracy
- Make a short pause between sentences to get better punctuation
Use Custom Vocabulary
Add the terms, the names, and the acronyms that you use often to your vocabulary list. This helps with:
- Proper nouns and company names
- Technical terms and industry terms
- Unusual spellings and brand names
- Common abbreviations that you want the app to expand
Choose the Right Model
- Cloud models — the accuracy is different for each provider. HyperWhisper Cloud’s ElevenLabs Scribe v2 tier has the highest accuracy rating, and Grok STT has a high accuracy rating. For the full comparison, see Choosing a Provider
- Local models keep your audio private and work offline, but they can be less accurate for specialized vocabulary
- A large local model usually gives better results, but it needs more processing time
Use a Dedicated Microphone
A dedicated microphone makes transcriptions much more accurate. A built-in laptop microphone picks up more background noise: keyboard sounds, fans, air conditioning, and room noise. This noise makes it more difficult for every transcription provider to isolate your voice.
More background noise in your recording makes it more difficult for the speech-to-text model to separate your words from the other sounds. This applies equally to cloud services and to local models.
These microphones are good options:
- USB microphones – Easy to install and good quality (for example, Blue Yeti, Audio-Technica ATR2100x)
- Headset microphones – These keep the same distance from your mouth
- Lavalier/lapel mics – Good for mobile use and when you move around
A low-cost USB microphone usually gives better accuracy than the built-in microphone of your computer.
Optimize Your Recording Environment
- Do a test of your microphone input level to make sure that HyperWhisper receives your voice clearly
- Decrease background noise: close the windows, switch off fans, and do not type during a recording
- To decrease echo, record in a carpeted room or a room with soft furniture
- Put your microphone 6–12 inches from your mouth to get the best signal-to-noise ratio