Guide
The mechanism holds. The percentage you read is folklore
Almost certainly yes, and the mechanism is not arguable. A large share of feed video is watched with the sound off. An uncaptioned talking-head ad says nothing at all to those viewers. What does not exist is a credible universal lift figure. This page covers where the famous statistic came from, the four jobs a caption does, and which caption decision is worth testing. By the end you will know what to stop arguing about.

- Jobs captions do, one of them access
- 4
- Caption packs shipped, no pack builder
- 6
- Credible universal lift figures in existence
- 0
A silent viewer reads nothing you said, and that is the whole argument
A large share of video in social feeds is watched without sound, particularly on first exposure and particularly in public. For those viewers an uncaptioned talking-head ad is a silent film of a stranger's face.
Whatever the true lift number is, it is the difference between communicating with those people and not communicating with them at all.
Be precise about that sentence, because everything below rests on it. Captions do not persuade anyone by existing. They make persuasion possible for the part of your audience that will never hear you.
That framing also tells you where the remaining upside sits. Once the words are on screen and legible, adding more captions does nothing.
The gains after that come from style, position, pacing and emphasis. That is a design question rather than an accessibility one.
So the recommendation is boring. Caption everything and stop debating whether to. Spend the argument you saved on which style, and on where it sits.
The famous statistic lost its qualifier somewhere around 2017
Most caption statistics trace back to platform marketing material or to one small study. Then they get relayed until the qualifier falls off.
A figure measured on autoplay preview behaviour in one placement in one year becomes a universal law about attention. It gets quoted to the decimal place by people who never saw the original.
Two distortions matter. The sound-off share varies enormously by platform, placement and context. A feed that autoplays muted produces different numbers from an app people open with headphones in.
The second is the outcome measured. Many quoted effects are on view-through or completion rather than on conversion. A caption pack can lift completion and leave sales flat.
None of that is a reason to skip captions. It is a reason to stop justifying them with borrowed numbers.
The argument from mechanism is stronger anyway. A colleague can dispute your percentage. Nobody can dispute that a silent viewer reads nothing you said.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Three of the four jobs are performance levers you are probably not pulling
Treating captions as a transcript undersells them. On a vertical ad they are a design element with several separate effects. Only the first is about access.
The other three change how long somebody stays. That is why caption style is worth testing and caption presence is not.
- Comprehension. The obvious one. Sound-off viewers get the message at all.
- Pace. Words appearing in rhythm create motion in an otherwise static shot, which holds a scrolling eye. A caption changing every beat does some of the job a cutaway would do.
- Emphasis. Making one word larger or a different colour is the closest thing video has to bold text. On the pivot word it directs attention. On every word it is noise.
- Repair. When somebody glances away for a second, the caption lets them rejoin mid-sentence. In a feed nobody rewinds, they leave.
What a caption is doing while it transcribes
Only the first is an accessibility job. The other three decide how long somebody stays.
Comprehension
Sound-off viewers get it at all
Pace
Motion in a static shot
Emphasis
One word, not every word
Repair
Rejoin after a glance away
Test which captions, not whether, and read your own name before you export
The clean test is boring and rarely run. The same cut, captioned and uncaptioned, live at the same time against the same audience. Most teams never run it because the answer feels obvious. They are probably right.
So the real question is a different one. Test which captions. Style, size, position and emphasis all vary. The differences show up in hold rate.
One test worth running is a heavy per-word emphasis pack against a plain readable style, on identical footage. The result is often larger than people expect, in both directions.
Set the verdict conditions before launch. Judge on a fast metric such as hold to the halfway point. Give each arm enough impressions to mean something.
Cutroom ships six caption and style packs. Clean-ugc, bold-impact, karaoke-flow, editorial, fast-native, minimal-lux. Colour, weight, size and position change on a cut that already exists. That makes this the cheapest test on the list to actually run.
One habit belongs here. Automatic transcription gets a product name wrong on the first pass. Every transcriber does, and read-back is not optional.
An ad that misspells your own brand in the frame where the name lands is worse than no ad. Fix it in the transcript, where the captions and the cut both read from the same text.
One word in twenty is the one that has to be right
Product names and unusual spellings are where transcription fails. Reading the transcript back costs a minute and protects the frame the name lands in.
Settle the caption argument once, then put the hour on the opening sentence
Captions work on viewers who already gave you two seconds. If the hook fails, the caption pack is decoration on a video nobody reached.
So when hold is fine and hook rate is poor, the afternoon belongs upstream. Rewrite the opening sentence. Reshoot the first frame. Those two carry the most variance in the whole ad.
Both are cheap to change here. Trim the opening. Delete a line and the cut rebuilds around the gap. Record twenty fresh seconds and run the take through again. Each version is a 10 credit export rather than an evening.
On the caption side, six packs ship. Colour, weight, size and position adjust inside each, plus per-word emphasis. There is no pack builder, so a signature look from another tool gets approximated rather than copied.
That combination is what makes captions stop being a decision. They arrive styled and timed with the cut. The hour you saved goes on the opening sentence, which is where it was always worth more.
Questions people ask
- Do I need captions if my ad has no speech?
- Usually yes, in the form of on-screen text rather than a transcript. A muted viewer with no speech and no text has nothing to read, so whatever the ad is claiming has to appear somewhere. Text on screen is doing the caption job under a different name.
- Are platform auto-captions good enough?
- They are convenient, inconsistent, and positioned wherever the app decides. If the same file goes to more than one destination, or if caption position and style are part of how the ad works, burn them in and keep control.
- Should every word be emphasised?
- No, and heavy emphasis on everything is the most common caption mistake. Emphasis works by contrast, so a pack that shouts every word carries no information. Mark the pivot word in a claim and leave the rest plain.
- Is a caption test worth running for everyone?
- If your ads already carry legible captions, the remaining gains are in style and emphasis and they are small. Spend the hour on the opening sentence instead. It has far more variance in outcome than any caption decision.
Caption everything and stop quoting the statistic. Here the captions arrive with the cut, styled and timed, so the only decision left is which of the six packs.