Feature
Highlight a phrase and the clip covers those exact words
B-roll is the part of an edit that eats the afternoon. Here the machine searches, picks the shots, and cuts each one to the words it covers. That happens inside the 100 credit batch, before you touch anything. Here is where the footage comes from, how much of the ad it may cover, and what a swap costs. By the end you will know how many shots land in a ninety second ad.

- Credits, search included
- 100
- Coverage ceiling at fast pace
- 52%
- Credits to swap a clip
- 0
It covers the concrete words and leaves the abstract ones alone
Before you touch anything, the machine has read your transcript. It hunts the words doing concrete work. A product, a number, a place, a result, a before and an after.
Those get covered. Abstractions mostly do not. A stock shot over the word therefore is noise with a licence fee.
It then searches for footage matching those words. Each clip is cut to the run of words it sits on. If a clip is shorter than the phrase, nothing is stretched or slowed. It ends, and you are back on camera.
Your own footage gets first pick. Your product, your premises, your faces, used wherever the words fit. Licensed public footage fills the rest.
When the machine picks wrong, and it will, swapping is one action and costs nothing. Replace it, or search millions of free clips for the shot you had in mind.
How much of the ad footage may cover
Pick the pace and you pick the ceiling. There is no setting that goes past 52 percent.
- Chill pace40%
- Normal pace45%
- Fast pace52%
Searching was never the expensive part. Placing was
Watch anybody place b-roll by hand and the ratio is the same. Two minutes to search. Twenty to place.
The twenty minutes are not craft. They are arithmetic done by a human with a mouse. They exist because the timeline thinks in seconds while you think in words.
Nobody has ever wanted a shot to begin at four seconds and thirty-one hundredths. They wanted it to begin on the word everything. So you translate an intention into a number, drag until it is close, then fix the translation error eight times an ad.
Multiply by eight shots, then by every variant you meant to test. That is where the month went. It was never a shortage of ideas.
Marking words removes the translation step. There is no number to get right, so there is no number to get wrong on the second version either.
Eight shots, placed by hand, for one ad
Searching was never the cost. Placing was, eight times, on every version you tried.
- Find one clip2 min
- Place that one clip20 min
- Eight shots in one ad160 min
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Coverage stops at 52 percent because the face is the trust
Coverage is capped by the pace you choose. 40 percent at chill, 45 at normal, 52 at fast. Even at the busiest setting you are on camera for close to half the ad.
People argue with that number before they test it, so here is the reasoning.
An ad that is 80 percent stock footage stops reading as a person talking to you. It starts reading as an ad. That is the moment the thumb moves. The face is the trust, and the footage shows what the face is talking about.
The cap also protects the through-line. Cut away too often and the viewer never settles back into who is speaking, so the argument never accumulates.
Once the ad is at its cap, a further highlight replaces a weaker clip rather than adding one. A ninety second ad at normal pace carries about forty seconds of footage, which is roughly eight shots.
Twenty seconds of an ad at normal pace
Nine seconds of footage, eleven of you. The cap is 45 percent and no highlight moves it.
It finds footage. It does not invent footage
It finds and places real clips. Yours first, then licensed public ones. If your line is about a product that does not exist yet, no prompt invents a shot of it. Upload your own frames instead.
It reads what you said, so what you said decides the results. A line like it changed everything gives the search nothing to grip. Name what changed and the clips improve at once.
Concrete words get good clips. Vague words get generic ones. That is a script problem wearing a footage costume, and it is a thirty second fix on the page.
Everything hangs on a talking-head take. The transcript is what the shots are placed against, so the read is the spine of the whole ad.
And coverage stops at the cap for the pace you picked. Half the ad is the ceiling by design, because the other half is the reason anybody believes it.
The search is inside the 100 credit batch, so an afternoon becomes a minute
Transcription, the directing pass and the footage search are one batch at 100 credits. That covers the searching, the choosing and the cutting-to-words for every clip in the ad.
Swapping a clip afterwards costs nothing. Export is 20 credits per minute of finished video, so a thirty second ad is about 10 credits to download.
The trial is 7 days and 300 credits with no card. Basic is $39.99 a month for 2,500 credits. Premium is $79.99 for 5,000. A $15 top-up adds 1,000 credits mid-month.
Compare that with the resource you were spending, which was never money. Eight shots at twenty minutes each is most of a working afternoon per ad.
An afternoon per ad is how an account ends up testing six creatives a month against a 5 to 8 percent hit rate. The credits were never the constraint. The afternoons were.
Questions people ask
- Where does the footage come from?
- Your own library first, so your product and your faces get used wherever they fit. Where there is no match, the gap is filled from licensed public footage.
- Can I choose the clip myself?
- Yes, and it costs nothing. Every spot can be swapped, and you can search millions of free clips and drop in the one you want. The machine's pick is a starting point.
- Why is a shot only covering part of my sentence?
- Clips are cut to the words they cover. If the clip is shorter than your phrase, it ends and you come back on camera rather than being stretched or frozen.
- Does it generate new video with AI?
- For b-roll, no. It searches and places existing footage. Generated presenter video is a separate module, where a photo and a voice-over become a lip-synced take.
- Is it right for every ad?
- It is built for ads where a person carries the read and footage shows what they are describing. Coverage tops out at 40, 45 or 52 percent by pace. If you want nobody on screen at all, the avatar module renders the presenter instead and the same footage goes underneath it.
Highlight the words that need showing. The shot lands on them, not near them, and the afternoon stays yours.