Feature

Instrumental for a voice-over: the empty box is the whole trick

The empty lyrics box is the whole trick. Leave it alone and the track comes back instrumental, for the same 20 credits and the same five lengths from 15 to 120 seconds. Getting one that sits under speech is a description problem, and the description has to name what should be missing. This page covers how to write it, why the level is almost never the fix, and how to brief against the read you already recorded.

Microphone and laptop setup for audio recording in a studio environment.
Photo by Jeremy Enns on Pexels
One instrumental, any length
20 CR
Where a speaking voice lives
500 Hz to 4 kHz
Editor controls: on, which track, how loud
3

How to get an instrumental, exactly

Leave the lyrics box empty. That is the only thing separating an instrumental from a song here.

Type the sound in the description box in your own words, comma separated. Ambient pads and texture, almost no rhythm is a complete brief.

Pick 15, 30, 60, 90 or 120 seconds. Match it to the ad rather than generating long and trimming.

Price is 20 credits flat at every length. The 300 free credits with no card cover fifteen attempts.

In the ad editor the music controls are turn it on, pick the track, and set how loud it sits.

There is no ducking, no stems and no tempo slider, so the description carries all of the control you have.

The instrumental in four steps

One of the four is leaving a box alone.

  1. Empty lyrics box

    This is what makes it instrumental

  2. Describe absences

    Sparse, no pads, nothing mid

  3. Match the length

    15, 30, 60, 90 or 120 seconds

  4. Generate and listen

    20 CR, on a phone

Describe what should be missing, because the voice needs the middle

A speaking voice occupies roughly 500 Hz to 4 kHz, and the consonants that carry meaning sit at the top of that band.

Anything in the track living in the same range masks those consonants, and the listener stops following without knowing why.

So a good brief names absences. Sparse. No pads. Nothing in the midrange. Low end and high percussion only. No vocals.

It reads oddly and it works. A brief asking for warmth, fullness or energy returns exactly the midrange you cannot afford.

Ask for slow, around 70 BPM, when the voice-over is fast, because two competing rhythms make both feel hurried.

The word to avoid is big. It reliably returns a wall of exactly the frequencies your voice needs.

  • Sparse, not full. No pads, not warm.
  • Low end and high percussion, nothing between
  • Slow the bed when the voice is fast, near 70 BPM
  • No vocals, ever, under speech

Who owns which part of the sound

The voice owns the middle. Everything you brief should live above or below it.

  • Under 500 Hz, the bed can have itBass, safe
  • 500 Hz to 4 kHz, the voice owns itKeep the bed out
  • Above 4 kHz, share it carefullyHigh percussion, safe

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

The level is almost never the fix

When a bed is fighting the voice, the instinct is to pull the music down until the problem goes away.

That works until the music is so quiet it contributes nothing, at which point you paid 20 credits for silence.

The actual problem is that the track is busy in the same band as the speech, and no level solves an overlap.

Generate a sparser track for another 20 credits. That fixes it in one step and leaves the music audible.

Judge it on a phone at half volume held at arm's length, not on headphones, because headphones separate what a phone cannot.

If the words are clear on a phone and the music is still noticeable, the description was right.

Match the bed to what the voice is doing, not to the brand

A fast, urgent read wants a slow, spacious bed. A slow, calm read can carry more movement underneath it.

That is the opposite of what most people brief, because they describe the brand rather than the delivery.

A bed is a supporting decision. It should be described against the voice-over that already exists, which means recording the voice first.

If you have the take, listen to it, then write the description while it plays in your head.

That order costs nothing and it produces a usable bed in one or two attempts rather than five.

Save the descriptions that worked. Three good sentences cover most of what a brand ever needs behind speech.

Reusing a description also gives you a family of beds rather than a set of unrelated ones, which is how a run of ads starts sounding like it came from the same place.

Three saved sentences carry a brand through a year of voice-overs

Three saved descriptions and a phone to check on will cover a year of ads. Nobody has to open an audio editor once.

What this does is narrow. It turns a sentence about a sound into a usable bed in one step, cheaply enough that a wrong answer costs almost nothing.

That converts a production problem into a writing problem. For anybody whose job is making ads rather than mixing them, that is the better problem.

Two ads want no bed at all. A dense voice-over with no pauses gains atmosphere and loses clarity, which is a bad trade on a direct response ad.

A demonstration where the product makes a real sound is stronger with that sound than with music over it.

Play the ad muted, then with voice only. If voice only holds, the bed is optional rather than required.

Four plain limits sit around all of this. No ducking, no stems, no editing after generation, and 120 seconds is the ceiling.

Frame-accurate audio work and automated ducking belong to a full audio editor. This module is not competing for that job.

The track is generated for you rather than licensed from a catalogue. Confirm the usage terms that apply to your account before running it behind paid spend.

Questions people ask

How do I make sure there are no vocals?
Leave the lyrics box empty, which is what produces an instrumental. Adding no vocals to the description as well does no harm and helps when the genre you asked for usually carries a vocal line.
The music is fighting my voice. What do I change first?
The description, not the level. Ask for sparse, nothing in the midrange, low end and high percussion only. A busy track cannot be fixed by turning it down, because the overlap is in frequency rather than volume.
Should I record the voice-over before generating the bed?
Yes. The bed should be briefed against the delivery that exists. A fast read wants a slow, spacious bed, and you cannot know the read is fast until you have heard it.
Who should not use a bed under a voice-over?
Anybody with a wall-to-wall read and no pauses. There is nowhere for music to sit and it will only cost clarity. Use silence, or generate a bed for the opening and closing seconds only.

Empty the lyrics box, describe the absences, and check it on a phone. That sequence produces a usable bed in one or two attempts.

Start with one take300 free credits · no card · cancel anytime