Guide
A bed level and a duck depth. One number cannot do both jobs
Set the music about 18 decibels under the mix, then drop it another 12 whenever somebody is speaking. That gives you a bed you can feel in the gaps and a voice that never fights anything, and it is two settings because a single fixed level is always either too loud under the read or inaudible between lines.
- Where the bed sits when nobody talks
- -18 dB
- Extra drop while the voice runs
- -12 dB
- Milliseconds in and out of the duck
- 200 / 400
One fixed level always fails at one end of the video
Pick a single music level and you are choosing which half of your video to ruin. Loud enough to be felt in the pauses is loud enough to muddy consonants under the read.
Quiet enough to stay out of the way under speech is quiet enough to vanish entirely in the gaps, at which point the music is doing nothing but adding file size.
So it is two numbers. The bed level is where music sits when nobody is talking. The duck is how much further it drops when somebody is.
We run the bed 18 decibels under the mix and duck a further 12, which puts music around 30 down while the voice runs. The general advice you will find elsewhere, 15 to 20 decibels below speech, describes the same relationship from the other side.
Where music sits, in each half of the video
The gap between the two rows is the duck. It is what makes music feel present without ever competing with a word.
- Bed, nobody speaking-18 dB
- Under speech-30 dB
- One sound effect-6 dB
The ramps matter more than the level, and almost nobody sets them
A duck that happens instantly is audible as an effect. You hear the music get shoved out of the way, which is worse than having it slightly too loud.
Two hundred milliseconds on the way down is slow enough to be invisible and fast enough to be out of the way before the first syllable lands. Four hundred on the way back up lets the bed return without swelling.
Asymmetric ramps are the trick. Fast in and slow out matches how attention works: you want the music gone the instant somebody speaks and you do not want it announcing itself when they stop.
Gaps shorter than the ramps do not get a duck at all. A quarter second of silence between two sentences is not an opportunity for the music to come up, it is a breath.
What the envelope does around one sentence
Four moves per sentence, none of which a listener should be able to name.
Speech starts
Ramp down over 200 ms
Voice runs
Bed holds at -30
Speech ends
Ramp up over 400 ms
Short gap
No duck at all, it is a breath
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
We tried a shallower duck first and it did not work
The first version dropped the bed by 8 decibels rather than 12. On paper that is plenty of separation and it sounded fine on monitors.
It was not fine. At 8 down the bed only fell to about 26 under the mix, which is close enough to the read to blur consonants, and the effect was worst on exactly the words that carry a claim.
Twelve puts music at 30 under speech. The voice wins cleanly, and because the bed still returns to 18 in the gaps, the track keeps its energy where there is room for it.
That is the whole argument for these being separate numbers. Deepening the duck cost nothing in the gaps, because the gaps are governed by the other setting.
When a ducked bed is the wrong shape entirely
If the music is the point, none of this applies. A track somebody is meant to listen to gets mixed for listening, and burying it 30 decibels under a voice defeats the exercise.
The same goes for anything rhythmic where the edit is cut to the music. There the voice serves the track, and ducking the track for the voice inverts the whole design.
Sound effects follow a different rule again. One at about 6 decibels under the mix reads as punctuation, and more than one every eight seconds or so reads as a cartoon.
Cutroom is built for the first case and not the second. It assumes a person talking with a bed underneath, so if you want music-led video, a timeline editor is the honest recommendation.
Questions people ask
- Is there one right number for music under a voice?
- No, there is a right relationship. Most published advice lands on music sitting 15 to 20 decibels below the speech, and the two settings above produce that. What varies is the track: a dense mix needs more separation than a sparse one.
- Can I just lower the music manually where the voice is?
- You can, and on a short video it is fine. It stops being fine around the fifth revision, because every script change means redrawing every ramp by hand.
- Does ducking hurt the music?
- It changes it. A bed that is ducked hard loses its sense of arrangement, so tracks with a lot of movement fare worse than simple ones. Choosing a simpler bed is usually better than ducking less.
- Does Cutroom let me set these myself?
- You can turn music on or off, choose the track and set how loud it sits. The duck depth and the ramp lengths are fixed, which is the right call for an ad and the wrong one if you want to mix it yourself.
Two settings, four numbers, and a listener who never notices any of them. That envelope is compiled into every Cutroom cut that has music on it, so the only decision left is whether you want a bed at all.