Listeners will forgive a little background hiss, a slightly dull tone or a voice that isn't quite radio-ready. What they won't forgive is having to rewind because they missed a word. For online tutorials, the voiceover is the instruction — everything else is supporting material. That reframes the whole job: you are not chasing a broadcast sound, you are chasing understanding on the first listen, often through laptop speakers, phone speakers or cheap earbuds.
The good news is that clarity comes from a handful of controllable factors — room reflections, microphone distance, speaking pace and plosive control — rather than expensive gear. Get those four roughly right and a modest microphone will outperform an excellent one used carelessly.
Most home recorders don't need acoustic panels on every wall. They need one decent corner. Hard, flat, parallel surfaces bounce your voice back into the microphone a few milliseconds late, and that smear is what makes home recordings sound boxy and tiring to follow.
Find a corner of a room with carpet or a rug, ideally away from a window. Stand or sit facing into the corner, with the microphone between you and the walls. Then soften every hard surface near your head:
You can hear the difference in seconds. Record ten seconds of speech, then clap once and listen back to the tail of the clap. If it rings, add more soft material. If it stops dead, you're in good shape.
The single biggest cause of uneven voiceover levels is a wandering mouth. Leaning in for emphasis and drifting back between sentences creates swings of six decibels or more, and no amount of compression fixes that gracefully.
Aim for a consistent 15 to 20 centimetres between your lips and the microphone capsule — roughly a hand span. Then hold it. Use a boom arm or a sturdy stand rather than holding the microphone, and mark the position with a piece of tape on the floor so you can return to it after a break.
Angle matters too. Speaking straight into a cardioid microphone exaggerates plosives and breath noise, so turn the microphone slightly off-axis, around 15 to 20 degrees, and aim the diaphragm at your mouth rather than your nose. Keep your script at eye level so you're not tilting down and projecting at the desk.
Set your input gain so normal speech peaks around -12 dBFS to -6 dBFS, with the loudest moment of the session still leaving headroom. Record at 24-bit, 48 kHz if your interface allows it. That's plenty of data for later analysis and processing, and the extra headroom means a sudden laugh or a dropped pen won't clip.
Tutorial narration works best somewhere between 130 and 155 words per minute. Faster than that and listeners lose the thread while they're looking at what you're showing them; slower and the delivery starts to sound patronising.
The trick isn't a constant rate — it's deliberate variation with generous gaps. Try this pattern:
Breathe from the diaphragm and take breaths at sentence boundaries rather than mid-clause. If you run out of air, the pitch rises and intelligibility drops at exactly the moment the sentence gets important.
Plosives — the bursts of air behind p, b and t — produce low-frequency thumps that survive even decent editing. A pop filter placed five to eight centimetres in front of the capsule, angled slightly, will catch most of them before they reach the diaphragm.
Reduce the rest at source rather than in the edit:
If a take goes wrong, don't stop mid-sentence. Clap once, pause, and repeat the whole sentence from the top. That clap creates a visible spike in the waveform, which makes finding the join trivial later.
Monitor on closed-back headphones while recording, then check the result on the worst speaker you own. A laptop speaker or phone is the honest test of whether your voiceover carries.
Before editing, glance at the waveform and any spectrogram view your software offers. You're looking for three things: a consistent average level, no flat-topped peaks, and dark blobs low in the frequency range that line up with plosives. That's audio data doing real work for you — it tells you whether to fix the performance or the processing.
Keep processing modest. A gentle high-pass filter around 80 Hz removes rumble without thinning the voice. Light compression at a 3:1 ratio with a slow attack and moderate release evens out the distance drift. Finish with loudness normalisation to roughly -16 LUFS integrated with a true peak ceiling of -1 dBTP, which sits comfortably across most platforms. Then listen all the way through, without watching the screen, and ask one question: did you understand every word the first time?
April 25, 2019 at 10:46 am
Take in the iconic skyline and visit the neighbourhood hangouts that you've only ever seen on TV. Take in the iconic skyline and visit the neighbourhood.
Want to be notified when we launch a new template or an udpate. Just sign up and we'll send you a notification by email.
Soldman Kell
April 25, 2019 at 10:46 am
Take in the iconic skyline and visit the neighbourhood hangouts that you've only ever seen on TV. Take in the iconic skyline and visit the neighbourhood.