Skip to content

Guide

Which caption style fits your video? Type a line and watch it play

The same sentence lands five different ways depending on how the captions move. The previewer on this page plays your actual line word by word in five styles — from the big one-word pop that Hormozi-era ads made standard to a calm lower-third — right in your browser, with nothing uploaded.

Caption previewer

Runs in your browser · no account needed

10 words · about 4.4s at normal · sample line

Caption style

Word speed

The type on this stage is whatever fonts your device ships — TikTok, Reels and every editor draw captions with their own. Read this preview for movement, size, pace and placement, never for exact pixels. And every feed app covers part of the bottom of the frame with its interface; the safe-zone checker shows exactly where.

One more honest difference: the clock here is steady, and real speech is not. In the Cutroom editor, captions are timed to your actual spoken words automatically, from the transcript of your take.

Caption styles you can play on this page
5
Caption packs that ship in the Cutroom editor
6
Words on screen at once in a two-word stack
2

The style is doing half the talking

A big share of vertical video plays with the sound off, and for those viewers the captions are not an accessory to the voice — they are the voice. The style they move in is the tone the video is heard in.

A single word slamming in at full width reads as urgency. A full line with one word lighting up reads as guidance. A small, steady line at the bottom reads as calm competence. Same sentence, three different speakers.

That is why previewing a style on somebody else's demo line tells you very little. The words carry the meaning and the style carries the delivery, and you can only judge the pair together.

The previewer above takes your own line and plays it word by word in five distinct styles, at three speeds, on a 9:16 stage — all in your browser, no account needed, nothing sent anywhere.

The table below is the shorthand: what each style signals, and the kind of video where that signal helps rather than fights.

Five caption styles, and the register each one speaks in

None of these is the best one. Each is a register, and it has to match what the line is actually saying.

What it signalsWhere it earns its place
Bold PopEnergy — one loud word at a timeHooks, claims, direct-response ads
KaraokeGuidance — read along with the voiceExplanations, how-tos, storytime
Clean Lower-ThirdCalm authority — the footage leadsFounder videos, demos, expert takes
Two-Word StackRhythm — punchy pairs, stackedLists, fast product points
Outline ShoutMeme energy — survives any backgroundBusy footage, reaction-style clips

The five styles, one by one

Bold Pop puts one word on screen at a time, scaled in big and centred. Nobody can skim it: the viewer reads at exactly the pace you speak, which is why the pattern dominates direct-response ads. It makes a quiet sentence feel loud — and a weak sentence feel louder, so the line has to hold.

Karaoke keeps the whole line on screen and lights the spoken word as it lands. The viewer can read ahead, which suits explanation: they always know where the sentence is going, and the moving highlight keeps them anchored to the voice.

The clean lower-third is the broadcast move: small, steady, at the bottom of the frame. It never competes with the footage, which makes it the right choice when the footage is the argument — a demo, a walkthrough, a founder talking straight into the lens.

The two-word stack shows punchy pairs, stacked and bold. It sits between the pop and the karaoke line: more rhythm than a full sentence, more context than a single word. Two words is also about what fits comfortably at a large size in a 9:16 frame.

Outline Shout is all caps with a heavy dark outline. The outline is not decoration — it is what keeps white text readable over a bright sky, a white shirt or a busy street, which is why the style owns reaction clips and street interviews.

Play your line in all five before committing. The mismatch between a style's register and what the sentence actually says is usually obvious within one run.

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

Hormozi captions: what the one-word pop actually does

Ask an editor for Hormozi captions and you are asking for the Bold Pop pattern: one word at a time, big, bold, centred, timed hard to the voice — the look Alex Hormozi's clips pushed into the mainstream of short-form advertising.

The mechanic underneath is pacing control. When only one word exists on screen, the viewer cannot skim, cannot read ahead, cannot half-watch. Their reading speed becomes your speaking speed for the whole video.

That grip is why the pattern converts on hooks and hard claims, and why it exhausts on longer explanations. Thirty seconds of one-word slams is a pitch; ninety seconds of them is a strobe light.

It is also a claim about energy, and the claim has to be true. A considered skincare brand shouting one word at a time is wearing a costume; the calm lower-third fits what that brand is actually saying.

Type your opening line above, switch to Bold Pop at the fast speed, and you have an honest Hormozi-caption preview of the sentence you would actually say — which settles the question faster than watching anybody else's example reel.

What a preview on this page can and cannot promise

The fonts here are your device's. TikTok, Instagram and every editor render captions with their own type, weights and smoothing, so treat what you see as an honest preview of movement, size, pace and placement — not of exact pixels. The previewer says this on its own face, because a tool that oversells its output is a tool you stop trusting.

Placement matters more than most people expect. Every feed app draws its interface over the bottom of the frame — the caption line, the audio attribution, the action buttons — and a caption parked in that zone gets covered. The checker at /tools/safe-zone-checker shows those zones on your own footage, also without an upload leaving the browser.

The timing here is a steady clock: every word gets the same slice of time at the speed you pick. Real speech is not steady — you lean on some words, rush past others, and leave pauses that carry meaning.

That gap is exactly what transcript timing closes. When captions are placed from the actual audio, each word appears the moment it is spoken and the emphasis lands where your voice put it.

So use this page for the decision it is good at: which register fits your line, at what size, at what pace. The clock-perfect version of that register comes from an edit timed to a real take.

From a previewed style to captions timed to your take

Cutroom starts where this previewer stops. Film one take of up to three minutes on your phone, upload it, and it comes back as a finished 9:16 video: cut, covered with footage, and captioned word by word from the transcript of your own audio.

Six caption packs ship in the editor — including a two-word DuoStack in the same family as the stack you just previewed — and switching the caption style, colour, weight, size and position is one control, not a re-edit.

Emphasis becomes a decision instead of an accident: mark the word your claim turns on and it pops in the captions. Fix a mis-heard word in the transcript and the caption is already right.

The pricing is one sum. A batch is 100 credits and the export is 20 credits per output minute, so a 60-second finished video comes to about 120 credits.

Three finished videos are free inside your first seven days, no card, on the 600 credits granted at signup, and they export with a Cutroom mark across the middle. Lite is $19.99 a month for 1,250 credits — about ten finished one-minute videos, more when they run shorter — and cancelling is one click.

Questions people ask

What is the best caption style for Reels?
There is no single best — a caption style is a register, and it has to match the energy of what you are saying. One-word pop styles suit hooks and hard claims, karaoke-style highlights suit explanation, and a clean lower-third suits demos and founder videos where the footage should lead. Play your own opening line in all five styles above; the mismatch is usually obvious within one run.
What are Hormozi-style captions?
One big bold word at a time, centred and timed tightly to the voice — the pattern Alex Hormozi's short-form clips made a default for direct-response ads. It works by locking the viewer's reading pace to the speaking pace, which grips on a thirty-second pitch and tires on a long explainer. The Bold Pop style in the previewer above is that pattern.
Do captions have to be timed word by word?
No, but word-level timing is what holds attention with the sound off, and it is the norm in short-form ads. Line-level captions are fine for calm formats like a lower-third. The honest catch is that word-level timing by hand is tedious, which is why transcript-driven editors place each caption on the word as it is actually spoken.
Will my captions look exactly like this preview on TikTok or Instagram?
No. This page renders with the fonts on your device, and every platform and editor draws captions with its own type. What the preview shows honestly is the movement, the pace, the size on a 9:16 frame and the placement — which is what you are actually choosing between when you choose a style.
Is my sentence uploaded when I use the previewer?
No. The previewer runs in your browser on this page, with no account needed. Nothing you type is sent or stored anywhere — close the tab and it is gone. If you type nothing, it plays a built-in sample line so you can still compare the styles.

Play the line in all five registers, keep the one that sounds like you — then film the take and let the edit time every word to your voice.

3 videos free, no card3 finished videos free in your first 7 days, no card. They carry a Cutroom mark; Lite at $19.99/month removes it