Guide
Burned-in captions are part of the picture itself, not a layer the player adds
Burned-in captions are caption text rendered permanently into the video's pixels, so they appear identically on every platform and cannot be switched off. The alternative is a sidecar file such as SRT, which the player overlays and the viewer can toggle. Ads burn captions in. Broadcast and streaming mostly do not.
- File that looks the same everywhere it plays
- 1
- Ways captions can exist: in the pixels or beside them
- 2
- Caption style packs shipped here
- 6
In the pixels or beside them, and everything follows from that
A caption can exist two ways. As text rendered into the frames during export, or as a timed text file the player draws on top at watch time.
Burned in, the words are pixels. Every device shows exactly what the editor saw, in the chosen font, position and rhythm, and nothing can move or disable them.
As a sidecar file, the words are data. The viewer can toggle them, screen readers and search can read them, translations can swap in, and every player styles them its own way.
Neither is better in general. They are different tools, and the choice is made by where the video will live and who controls the player.
Ads burn in because the feed strips everything else
Vertical feeds give an advertiser no reliable caption layer. Platform auto-captions vary in accuracy, sit where the app decides, and look different on every placement.
Meanwhile most feed viewing starts muted, so the captions are not an accessory. For a large share of viewers they are the ad.
Burning in is the only way to control what those viewers experience. The words appear exactly where the safe zones allow, styled to be readable, timed to the speech, identical on every platform the file reaches.
Style is also doing persuasion work in this format. Word-by-word timing creates motion in a static shot, and emphasis on the one word carrying the claim is the video equivalent of bold text. None of that survives a plain text sidecar.
Burned-in against sidecar captions
The honest trade. Ads live in the left column because feeds offer nothing reliable on the right.
| Burned in | Sidecar file | |
|---|---|---|
| Identical everywhere | Yes, they are pixels | No, each player restyles |
| Viewer can toggle | Never | Yes |
| Styling and emphasis | Full control | Player decides |
| Screen reader access | No, it is an image | Yes, it is text |
| Fixing a typo | Re-export the video | Edit one text file |
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
The costs are real, and one of them bites on export day
Permanence cuts both ways. A transcription error in burned-in captions ships inside the file, and the fix is a re-export rather than a text edit.
That makes the transcript read-back the critical habit. Product names, prices and unusual spellings are exactly where speech-to-text misses, and exactly the words that must not be wrong on screen.
Accessibility is the second cost. Burned-in text is invisible to screen readers, which is one reason platforms and broadcasters prefer real caption tracks. Where a destination accepts a caption file, providing one alongside is the more accessible choice.
Language is the third. Burned-in captions are one language forever. A sidecar workflow can carry translations. A burned-in workflow re-exports per language or does not translate at all.
In practice: style once, check the name, then stop deciding
A caption style is a brand asset that should not be redesigned per video. Pick a readable style, position it inside the safe zones, and reuse it so the library ages as one body of work.
Cutroom burns captions in from the transcript, styled by one of six shipped packs with colour, weight, size and position adjustable inside each. There is no SRT export, which is the honest description of an ad tool rather than a subtitling suite.
Emphasis is a decision per claim, not per word. One emphasised word in a sentence directs the eye. Every word emphasised is noise in a bold font.
Then the boring habit. Read the transcript before export, fix the mis-heard word there, and the captions and the cut both inherit the correction.
Questions people ask
- Are burned-in captions the same as open captions?
- Yes. Open captions is the broadcast term for captions that are always visible because they are part of the picture. Closed captions are the toggleable kind delivered as data. Burned in and open describe the same thing.
- Do burned-in captions affect video quality or file size?
- Marginally. The text becomes part of the encoded image, so file size changes little. Sharp caption edges can shimmer on heavy compression, which is one more reason to use a clean, adequately sized style.
- Can I get an SRT file out of Cutroom?
- No. Captions here exist to be burned into the export, styled and timed for feeds. If a destination requires a separate caption file, generate one with a subtitling tool from the finished video.
- Should YouTube videos burn captions in?
- Long-form YouTube generally should not, because the platform has a real caption system viewers control and search reads. Shorts behave like a feed, where burned-in styling does the same work it does everywhere else vertical.
Pixels where the platform offers nothing better, data where it does. For feed ads that answer is pixels, and Cutroom styles and times them from the transcript in the same pass as the cut.