Automatic transcription has become good enough that, for most recordings, you read the subtitles once and fix a few words. The gap between "a few words" and "every other line" is almost always the audio, not the software. These are the habits that make the biggest difference.
1. Get the microphone close
Distance is the number one cause of bad transcripts. A phone on the desk two metres away picks up the room, the fridge and the echo of your own voice. A lapel mic, a headset or even earbuds with a built-in microphone, 10 to 30 cm from your mouth, removes most of that. The recording sounds "dry", and dry speech transcribes well.
2. Record in a soft room
Hard surfaces reflect sound and smear consonants together. Carpets, curtains, a sofa, a bookshelf, even a blanket on the desk, absorb the reflections. If the room rings when you clap, it will hurt the transcript.
3. Keep background music low, or add it later
Speech recognition separates voice from noise surprisingly well, but music with vocals is the hardest case because it is also speech. If you add a soundtrack, do it after transcription, or keep the music at least 15 to 20 dB below the voice. For the best result, export a voice-only track for the subtitles and the mixed track for the video.
4. One speaker at a time
Crosstalk, where two people speak at once, cannot be transcribed correctly by anyone, human or machine. In interviews and podcasts, let each speaker finish. If you record each person on a separate microphone, the mix will be cleaner and the transcript better.
5. Say names and acronyms clearly, once
Rare proper nouns, product names and acronyms are the words most likely to come out wrong. Say them clearly the first time they appear. After transcription, search the SRT file for the name and fix every occurrence in a text editor with find and replace. It takes a minute.
6. Export with sensible audio settings
The transcription does not need studio settings, but it does need a clean export:
- Sample rate 44.1 or 48 kHz, 16-bit, mono or stereo. Anything standard is fine.
- Avoid heavy noise reduction or aggressive compression plugins before transcription. They can smear the very consonants the recogniser listens for.
- Do not normalise to the point of clipping. Distorted peaks become wrong words.
- If the video file is large, export the audio alone (M4A, MP3 at 128 kbps or higher, WAV). Transcription uses only the audio, so the result is identical and the upload is faster.
7. Transcribe the final cut
Generate subtitles from the finished video, after trimming and re-timing. A subtitle file made from an earlier edit will be out of sync with the final one. If you must re-edit, generate again rather than shifting cues by hand. See how to fix out-of-sync subtitles for when that is unavoidable.
What to check after transcription
Even with perfect audio, read the subtitles through once. The places to look:
- Names and brands. Find and replace.
- Numbers and dates. "Twenty twenty-six" versus "2026" is a style choice; pick one.
- Homophones. "There" and "their", "its" and "it's". Rare, but they slip through.
- Line length. If a cue is too long to read, split it in a subtitle editor.
Get subtitles from your recording
Upload the video or audio file to ToneLark (up to 2 GB). The spoken language is detected automatically, and you get an email when the SRT and VTT files are ready. Open the SRT in any text editor for the checks above, then use it wherever you need it: on YouTube, in your video editor, or burned into the picture for social media.