Skip to content

Feature

Music under a voice-over works when it avoids one frequency band

A speaking voice occupies roughly 500 Hz to 4 kHz, and a bed works when it stays out of that band. Mood, genre and tempo are downstream of that one fact. This page covers why masking is not about volume, what to type to get a bed that survives, and how a phone speaker changes the answer.

A young woman wearing headphones sits thoughtfully at a table indoors.
Photo by Julio Lopez on Pexels
The band speech occupies
500 Hz to 4 kHz
One instrumental, any of five lengths
50 CR
A safe tempo under a fast read
70 BPM

Masking is the mechanism, and it is not about volume

Two sounds in the same frequency band compete. The louder one hides the quieter one. That is masking, and it is why turning the music down often does not help.

Consonants are the quietest part of speech and they sit at the top of the speech band. They are the first casualty.

Lose the consonants and the listener still hears a voice. They stop following the words, usually without knowing why.

That is the state most muddy ads are in. The complaint is the music is too loud, and the cause is the music is in the wrong place.

A quiet bed packed with midrange masks more than a louder bed that has nothing in the middle.

So the fix is a different track, not a different level, and a different track costs 50 credits.

Which parts of a track compete with speech

Everything the bed puts between 500 Hz and 4 kHz is taken directly out of your voice-over's clarity.

  • Sub bass and bass, under 500 HzSafe
  • Pads and keys, 500 Hz to 2 kHzWorst offender
  • Guitars and synth leads, 1 to 4 kHzEats consonants
  • Hats and shakers, above 6 kHzMostly safe

Describe the absences, because that is the only control you have

The module gives you one text box for the sound and one for lyrics. Leave the lyrics box empty for an instrumental.

In the description, name what should not be there. Sparse. No pads. Nothing in the midrange. Low end and high percussion only.

That reads like an odd way to brief music and it returns a far more usable bed than any request for warmth or energy.

Words to avoid: big, full, warm, lush, epic. Every one of them returns midrange.

Words that help: sparse, minimal, low end, high percussion, almost no rhythm, slow.

In the ad editor the controls are turn music on, pick the track, and set how loud it sits. No ducking and no stems, so the description is the mix.

  • Say sparse, minimal, low end and high percussion
  • Never say big, full, warm, lush or epic
  • Leave the lyrics box empty. A sung line is a second voice.
  • Ask for slow, around 70 BPM, under a fast read

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

Tempo has to disagree with the read, not match it

A fast voice-over over a fast bed makes both feel hurried, and neither one has anywhere to land.

A fast read over a slow, spacious bed reads as urgent and controlled, which is the tone most direct response ads are aiming at.

So brief the tempo against the delivery rather than against the brand.

Which means recording the voice-over first, listening to it, and writing the music description afterwards.

Most people do this backwards, generate the bed from a brand mood board, and then wonder why it fights.

Around 70 BPM is a reliable starting point under speech, and around 140 works when there is no voice at all.

The order that produces a usable bed

The bed is a response to the read. Generating it first is the common mistake.

  1. Record the voice-over

    First, always

  2. Listen to the pace

    Fast read wants a slow bed

  3. Describe absences

    Sparse, no midrange, no vocals

  4. Generate

    50 CR, one attempt usually enough

  5. Check on a phone

    Half volume, arm's length

A phone speaker removes the safe half of your track

A phone speaker reproduces almost nothing below roughly 500 Hz, which is exactly the band you were briefing the bed into.

So on the device most ads are watched on, the safe part of your track is inaudible and only the competing part survives.

That is why a mix that felt balanced on headphones turns muddy on a phone, and it is not a skill problem.

Judge every ad on a phone at half volume held at arm's length. Headphones will lie to you about this specific thing.

If the words go unclear on a phone, generate a sparser track rather than lowering the level to the point of pointlessness.

High percussion is the part of a bed that survives a phone speaker without touching speech. That is why it is worth naming in the description.

Name it explicitly. Shakers, hats, a light tambourine. Those three words do more for a phone mix than any level adjustment.

One rule removes most of the trial and error

The muddiness people blame on volume is nearly always a frequency problem. That is the whole finding.

Once you know the voice owns 500 Hz to 4 kHz, the brief writes itself. Most of the trial and error disappears.

Here that rule is the entire control surface, because the description is the mix. One sentence, 50 credits, one attempt.

Two ads want no bed at all. A voice-over with no pauses has nowhere for music to sit, and adding one only costs clarity.

A demonstration where the product makes a real sound is stronger with that sound than with a track over it.

Play the ad muted, then voice only. If voice only holds, the music is a decision rather than a requirement.

Four plain limits sit around it. No ducking, no stems, no editing after generation, and 120 seconds is the ceiling on a track.

Automated ducking and a frame-accurate mix belong to a full audio editor. This module does not compete for that work.

The track is generated for you rather than licensed from a catalogue. Confirm the usage terms that apply to your own account before running it behind paid spend.

Questions people ask

Why does turning the music down not fix the muddiness?
Because the problem is frequency overlap rather than volume. A quiet track full of midrange still masks the consonants in your voice-over. Generating a sparser track for 50 credits fixes it in one step and keeps the music audible.
What tempo works under speech?
Slower than the read. Around 70 BPM is a reliable starting point under a fast voice-over. Matching the tempo of the delivery makes both feel rushed, because two rhythms are competing for the same sense of pace.
Which instruments are safest under a voice?
Bass below 500 Hz and high percussion above 6 kHz. Pads, keys, guitars and synth leads all sit in the speech band and are the usual cause of an ad sounding muddy on a phone.
Who should skip the bed entirely?
Anybody with a wall-to-wall read and no pauses, and anybody whose product makes a sound worth hearing. Music helps in the gaps. An ad with no gaps has nowhere to put it, and adding one only costs you words.

Keep the bed out of 500 Hz to 4 kHz. That one rule solves more mixing complaints than every level adjustment combined.

3 videos free, no card3 finished videos free in your first 7 days, no card. They carry a Cutroom mark; Lite at $19.99/month removes it