Guide
Two words describe a thousand pictures. The query is the problem
Automatic cutaway selection goes wrong for a boring reason: the search term pulled out of your sentence is usually one or two words, and a stock library holds thousands of pictures that match it. What separates a tool you can use from one you cannot is not how clever the matching is, it is what happens when the match is weak.
- Words in a typical extracted query
- 1-2
- Clips a good pipeline places when unsure
- 0
- Honest budget for swapping the misses
- 2 min
The query is short because the sentence was short, and short queries are ambiguous
A sentence like we run the numbers before we quote yields a query like numbers. That word describes a calculator, a spreadsheet, a roulette wheel, a house door and a child's building blocks.
The tool did not misunderstand you. It got exactly what it asked for, and what it asked for was ambiguous.
This is why the failures look so strange. They are not near misses, they are direct hits on the other meaning of a word, which reads as much more broken than a vague result would.
Longer queries help and are not free. Ask for a woman in her thirties reviewing a spreadsheet in a bright office and you will match nothing at all, because that clip does not exist in any library.
How a sentence becomes a wrong picture
Nothing here is broken. Each step does its job and the ambiguity survives all of them.
Your sentence
Fifteen words of context
The extracted query
One or two of them
The library
Thousands of matches
The pick
Right word, wrong meaning
Stock is indexed by whatever the uploader typed, not by what is in the picture
Library metadata is written by people who want their clip to be found, so the tags are broad and hopeful rather than accurate. A clip tagged business, meeting, teamwork, success and growth is telling you nothing about what is on screen.
That is the raw material every automatic tool works from, ours included. Better matching on top of vague labels still gives you vague labels.
It also explains a specific failure people notice: the same handful of clips appearing across unrelated videos. Those are the ones tagged hardest, so they surface for everything.
The practical consequence is that the quality of your cutaways is bounded by the library long before it is bounded by the software.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
What separates a usable tool is the floor, not the matching
The important question is what happens when nothing good is found. A tool that always fills the slot will always put something wrong on your screen.
Our answer is a floor. If a candidate is rejected both by the words and by the model that scores the shortlist, it is thrown out rather than ranked, and the slot stays empty.
An empty slot is not a failure. It means the viewer stays on the speaker for another two seconds, which is a completely acceptable outcome and is often better than the clip would have been.
The other half is that your own uploads outrank stock. Footage you shot of your actual product is specific in a way no library is, and a pipeline that ignores it in favour of a polished stranger is optimising for the wrong thing.
Two ways to handle a weak match
The difference shows up on exactly the slots where the query was ambiguous, which is most of them.
| Always fill the slot | Refuse below a floor | |
|---|---|---|
| When the match is good | Correct clip | Correct clip |
| When the match is weak | Something wrong on screen | Stays on the speaker |
| What you fix afterwards | Hunt for the bad ones | Add where you want more |
| How it reads to a viewer | Assembled by a machine | Edited |
Plan on swapping two of them, and pick a tool where that is cheap
No automatic selection gets every slot right, and any tool claiming otherwise is describing a demo rather than a week of work.
The number that matters is how long a swap takes. If fixing one cutaway means re-rendering the video or paying again, the automation has not saved you anything.
In our editor you click the clip you disagree with and pick another, or search the free libraries yourself, and the cut rebuilds around it. Highlighting a phrase places a clip over exactly those words.
Budget two minutes per finished ad for this. That is the honest cost, and any pitch that leaves it out is selling you the demo.
Questions people ask
- Why does automatic b-roll keep using the same few clips?
- Because library metadata is written for discovery. The clips tagged most aggressively surface for the widest range of queries, so they appear everywhere.
- Would longer, more descriptive queries fix it?
- Partly, and they introduce the opposite failure. Very specific queries match nothing, and a tool that finds nothing either leaves the slot empty or falls back to something generic anyway.
- Is generated footage better than stock for this?
- Not for anything a customer recognises. Generated video is convincing on generic subjects and unreliable on your specific product, where the packaging shifts between shots and the logo becomes a smudge.
- How many cutaways should I expect to change?
- On a thirty second ad, expect to swap one or two. If you are changing most of them, the script is probably too abstract for any library to illustrate.
Short queries are ambiguous, libraries are hopefully tagged, and the honest fix is a floor plus a fast swap. That is why an empty slot is an allowed outcome here and changing a clip takes one click.