← Notes

Recording and placing a voice-over

The room matters more than the microphone, and the edit matters more than both. Here is how to get a usable voice-over on a Mac without buying anything.

A voice-over rescues footage that does not explain itself: a screen recording, a silent process shot, a sequence that needs context. It is also the fastest way to make a competent video sound amateur, and the reasons are mostly not about equipment.

The room beats the microphone

A £400 microphone in a bare room with hard walls sounds worse than a laptop microphone in a room full of soft furnishings. Reflections are what make recordings sound cheap, and no processing removes them convincingly.

What actually helps, in order of effect:

  • Soft surfaces. Curtains, a rug, a sofa, a bed. A bedroom with the curtains drawn is one of the better rooms in most homes.
  • Avoid the corner. Corners and parallel bare walls are the worst places. The middle of a furnished room is better.
  • Get close. Fifteen to twenty centimetres from the microphone. The closer you are, the more of what it picks up is you rather than the room. This is the single biggest improvement available for free.
  • Slightly off-axis. Speak past the microphone rather than directly into it, which reduces plosives — the thumps on p and b.
  • Kill the noise floor. Fridge, fan, air conditioning, laptop fans. Recording a heavy export while narrating over it is a classic mistake.

The much-repeated advice about recording in a wardrobe is genuinely sound. Clothes absorb reflections better than anything else most people own.

Levels

Record so that normal speech peaks somewhere around −12 to −6 dB, leaving headroom. Recording too hot and clipping is unrecoverable; recording too quietly and raising it afterwards brings the room noise up with it.

Do a thirty-second test at your normal speaking volume — not a test voice, which is always quieter than the real thing — and check before committing to a take.

Performance, which is most of it

Write it down, then read it as though you had not. Scripted delivery is the difference between a tight ninety seconds and a rambling four minutes. Reading it stiffly is the risk; reading aloud twice before recording fixes most of that.

Stand up. It changes your breathing and your energy audibly.

Go slightly slower than feels natural. Almost everyone speeds up when recording. What feels laboured in the moment usually sounds normal on playback.

Do not restart from the top. When you fumble a line, pause, breathe, and say the line again from the start of the sentence. You will cut the bad take out later. Restarting from the beginning every time is how a two-minute script takes an hour.

Record thirty seconds of silence. Just the room, no speaking. This is your noise profile if you need to reduce hiss later, and it costs nothing to have.

Editing it

Cut in the gaps between sentences, with a few frames of room tone at each end. Cutting hard against a word clips the consonant.

Leave the breaths in. Removing every breath sounds unnatural and slightly unsettling. Remove the loud ones, keep the rest.

Trim the long pauses. The gaps that feel right while speaking are usually too long on playback.

Duck the music. If there is a music bed, it needs to drop by several decibels while the voice is present. Music at a constant level either drowns the voice or is inaudible — there is no setting that works for both.

Match the voice to the picture, not the reverse. Write to the footage you have, then adjust clip timing to fit the narration where needed.

Placing it against the video

The sequence that works: assemble the picture roughly, record the voice-over against that assembly, then adjust the picture to the voice. Trying to record narration to a locked edit means fighting for timing on every line.

Fades at the start and end of the voice track prevent the small click that comes from a waveform starting abruptly.

Doing this in Garfi

Garfi Video Editor handles voice-over on the same timeline as the picture, so recording against the assembly and then trimming both together is one document rather than two applications.

You get audio waveforms — which is how you find the sentence boundaries by eye rather than scrubbing — plus the ability to detach audio from a clip, work with clip audio, music and voice-over as separate elements, and add fades where a cut needs smoothing.

Editable automatic captions are on the same timeline, which is worth pairing with a voice-over since most people watch muted first. The captions are a draft to correct, not output.

Export is MOV or MP4 in H.264 or HEVC, up to 4K where supported, with no watermark. Projects save as portable .gvideo packages, so re-recording one line next week means opening a file rather than rebuilding.

Free download, an export trial you start yourself — fourteen days, no automatic renewal — and a one-time purchase to unlock export permanently. No account, no cloud backend, no generative AI.

Keep reading