Feature
Karaoke captions: the sweep tells them where they are
Every other caption pack reveals a word as it is spoken. This one shows the whole line and lights the word being said. The eye reads ahead and never loses its place. That is worth a lot on anything with steps in it. This page covers what lands on screen, why mid-frame is deliberate, and which scripts it suits. By the end you will know whether your script has an order or a punchline.

- Words readable at once, one highlighted
- 4
- Runtime covered with footage at normal pace
- 45%
- Where the line sits in the frame
- Mid
What lands on screen
A full line of up to four words, positioned in the middle of the frame rather than low.
A highlight sweeps across the word currently being spoken, so the line ahead is readable and the place is never lost.
That is the mechanical difference from every other pack here: the words arrive before they are said rather than with them.
Transitions are crossfades plus a slide-up push, which is a clean explainer move rather than an aggressive one.
Zoom is moderate, b-roll arrives fullscreen, and the pace is normal, capping coverage at 45 percent.
Sound effects are on, which suits a piece with clear steps because each step gets a small marker.
- Caption engine: full line visible, current word swept
- Up to 4 words per group, mid-frame position
- Transitions: crossfade and slide-up
- Zoom: moderate. Pace: normal, coverage caps at 45 percent
- Sound effects: on
Why reading ahead changes comprehension
Every other pack reveals words as they are spoken. This one shows the line and marks the position.
The line appears
Up to four words, all readable
The sweep starts
On the word being spoken
The eye reads ahead
And knows what is coming
The line turns over
No place lost
It is for sequences, not for claims
An explainer has an order. First this, then that, and here is why the second part matters.
A viewer following an order needs to know where they are, and a sweeping highlight is the cheapest way to tell them.
That is different from a claim, which needs impact rather than orientation.
So this pack outperforms on how-to content, comparisons, process walkthroughs and anything with numbered steps.
It underperforms on a single strong offer, where reading ahead removes the punch of a line arriving whole.
Sort your script by whether it has an order or a claim. The pack choice then makes itself.
The gain is easiest to see on a numbered walkthrough. A viewer who can read the next step while hearing the current one stops rewinding.
That matters most on the steps nobody expects. A surprising instruction is the one a viewer misses when the words arrive at the speed of speech.
A useful signal is whether the word then appears in your script. Scripts with then in them have an order and belong here.
Scripts built on because and instead are argument rather than sequence. Those do better in a pack that lands lines whole.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Mid-frame position is a decision, not a default
Most caption packs sit the words low, out of the way of a face.
This one puts them in the middle, where the eye already is, because reading is the primary activity rather than a secondary one.
That covers more of the picture, which is a cost worth paying when the words carry the content and the picture is illustration.
It is the wrong trade when the face is the content, which is why clean-ugc keeps its line low.
You can change the position in the editor along with colour, weight and size. None of that re-cuts the ad.
Moving the line down loses part of what makes this pack work. Try it as drafted first.
How much of the frame the captions occupy
Mid-frame placement covers more picture. On an explainer the words are the content, so the trade is worth it.
- minimal-lux, low and smallBarely there
- clean-ugc, low single lineOut of the way
- editorial, lower thirdFramed
- karaoke-flow, mid frameIn the eyeline
- bold-impact, stacked capsDominant
Four ads that want a different pack, and the hybrid that keeps both
A price-led offer with a deadline. Reading ahead removes surprise. Surprise is what an offer line runs on.
A testimonial, where mid-frame text sits over the face you filmed for the purpose of being seen.
A luxury or beauty product, where a sweeping highlight reads as instructional rather than desirable.
A montage with no continuous speech, since there is no sentence to keep a place in.
All five alternatives sit in the same batch. Run the cut through karaoke-flow and clean-ugc and compare on a phone.
The right answer is usually obvious in the first four seconds of each.
There is also a hybrid worth knowing. An explainer that ends on an offer can run in karaoke-flow throughout.
Emphasise a word so it pops in the captions. The closing line then gets the weight a louder pack would have given the whole ad.
That keeps the comprehension advantage through the body and buys back the impact at the end.
It costs nothing to try. Switching packs and adding emphasis are both inside the batch you already paid for.
One thing to avoid. Emphasising several words in the same ad spreads a scarce resource thin and removes the effect entirely.
On a genuine list, the sweep beats every other decision here
If your explainer is genuinely a sequence, this pack does more for comprehension than anything else available in the editor.
It arrives with the batch at 100 credits. The export is 20 credits per output minute.
Four plain facts sit alongside it. No logo and no fixed closing frame lands on the export. Nothing of yours is added to the file.
There is no timeline, no keyframes and no layers. Position, colour, weight and size are settings.
Every export is 9:16.
And the sweep is orientation. Orientation only helps when there is somewhere to be, so a script with no order belongs in another pack.
Questions people ask
- How is this different from word-by-word captions?
- Word-by-word reveals each word as it is spoken. This shows the whole group of up to four words and sweeps a highlight across the one being said, so the eye can read ahead. That difference is what helps comprehension on a sequence.
- Why are the captions in the middle of the frame?
- Because on an explainer the words are the content rather than a subtitle over the content. Mid-frame is where the eye already is. You can move them down in the editor, though it removes part of what makes this pack work.
- Does it work without a voice-over?
- Poorly. The sweep tracks speech, so a montage with no continuous talking has nothing to track. Use bold-impact or fast-native for montage-led ads.
- Who should not use karaoke captions?
- Anybody running a single strong offer. Letting a viewer read ahead removes the impact of a line landing whole, which is exactly what an offer needs. Use bold-impact for those and keep this pack for anything with steps.
Four words visible, one word lit. No other pack tells a viewer where they are in a sentence, and it arrives with the batch.