← Notes

Music under dialogue without drowning it

The music sounded right while editing and buries the voice on a phone speaker. Here is how loud music should sit under speech, why some tracks fight dialogue no matter what, and a mixing routine that works on small speakers.

You add music under a talking-head video, set a level that sounds fine through your headphones, and export. On a phone speaker, the music swallows every other word. Or you turn it down so far that it disappears entirely except in the pauses.

Getting music to sit under speech is a matter of level, choice of track and a little movement. Nothing here needs special equipment.

How much quieter

Speech needs to be clearly on top. As a rule of thumb, music under dialogue sits somewhere around 15 to 25 decibels below the voice — much quieter than feels natural while editing.

Why so much? Because you already know what is being said. You wrote it or recorded it, so your brain fills in words the music covers. A viewer hearing it for the first time cannot.

A practical test: play it back quietly, at the level someone might watch on a phone in a room with other people. If you have to strain to understand a sentence, the music is too loud. Then play it back on the worst speaker you own — a laptop or phone speaker. That is where most videos are watched.

Waveforms tell you a lot

The editor's waveforms show loudness visually. Speech should have tall, spiky waveforms. Music underneath should be a noticeably lower, flatter band. If the music waveform is as tall as the speech waveform, it is too loud, regardless of how it sounds in your headphones.

Ducking: louder in the gaps

Constant low music can feel flat in pauses. The professional answer is ducking: the music is lower while someone speaks, and rises a little in the pauses, intros and outros.

Automatic ducking tools do this for you. By hand, the approach is:

  1. Split the music track where speech starts and stops.
  2. Set the sections under speech lower, and the gaps higher.
  3. Add short fades at each change — a fraction of a second to a second — so the level moves smoothly rather than jumping.

You do not need to follow every breath. Duck for sentences and paragraphs, not individual words.

Choosing music that does not fight

Some music is impossible to put under speech at any level:

Vocals. Lyrics compete directly with dialogue, because the brain tries to follow both. Use instrumental tracks — many libraries offer instrumental versions.

Busy mid-range. Human speech sits mainly in the middle of the frequency range. Music with lots of activity there — lead guitars, piano melodies, saxophones — masks the voice. Music whose energy is in the low bass and high shimmer, with a quieter middle, leaves room.

Strong melody and big changes. A memorable tune pulls attention. For background use, simpler, steadier, repetitive tracks work better. Save the melody for the intro, outro and moments without speech.

Tempo that fights the speaker. Very fast music under a calm explanation feels anxious. Match energy to the content.

Loudness for the whole video

Once dialogue and music are balanced against each other, the overall level matters too. Platforms normalise loudness to a target, turning loud videos down and sometimes quiet ones up. A video mixed very loud gains nothing; one mixed very quiet may sound weak next to others.

Aim for dialogue that peaks comfortably without clipping — the waveform never flattening at the top — and let the platform handle the rest.

Rights

Music is almost always copyrighted. Using a track you do not have rights to can get a video muted, blocked or demonetised, even if it is only a few seconds.

Safe sources are libraries that license music for video use, platform-provided audio libraries for use on that platform, and music you made or commissioned. A licence to listen to music on a streaming service does not include using it in your videos.

A mixing routine

  1. Set the dialogue level first. Consistent, clear, not clipping.
  2. Add the music and set it low — lower than you think.
  3. Duck under speech, with short fades at each change.
  4. Fade in at the start and out at the end. A music track that stops abruptly mid-phrase sounds like a mistake.
  5. End music on a natural phrase where possible: trim so the video ends where the music does.
  6. Check on a phone speaker, at low volume.

What Garfi does

Garfi Video Editor is a native Mac editor, and the routine above uses its audio tools directly. Clip audio, music and voice-over sit on the timeline with waveforms, so you can see whether music is sitting low enough. Split a music track where speech starts and stops, set each section's level, and add fades so the changes are smooth. Detach audio separates a clip's sound from its picture when you need to treat them differently.

Editable automatic captions, using Apple speech recognition, are a safety net for viewers watching with the sound off — or in a noisy room where even a well-balanced mix is hard to hear.

Export is MOV or MP4 in H.264 or HEVC, up to 4K when supported, with no watermark. The download is free with an export trial you start yourself; a one-time purchase unlocks export permanently. There is no account, developer cloud or generative AI.

Keep reading