Guide
The output is not badly written. It is generically written.
Generated ads underperform, and video quality has nothing to do with it. Three named failures do. The script is the middle of everything ever written about your category. Nobody is standing behind the claim. Within months the whole category converges on one look. This page covers each failure in turn. Then the ten-second test that catches the first one. Then the line between work worth automating and work worth keeping.

- Named failures, and none of them is video quality
- 3
- Specifics about your business a model can know
- 0
- Test: delete your brand name and read it back
- 1
Failure one: it returns the middle of everything written about your category
That is what a language model is for. It is also why the output feels familiar, and familiar is fatal in a feed.
Your viewer has scrolled past a thousand versions of 'tired of wasting hours on this'. The thumb moves before the brain engages.
The generated script is not badly written. It is generically written. That is worse, because bad writing at least surprises.
The fix is not a better prompt. It is supplying the specifics a model cannot know. The actual price. The objection from your inbox. The sentence a customer sent you last Tuesday.
Prompting harder does not fix it. A longer brief moves the output towards the average of a narrower category. That is still an average.
Test any script this way. Delete your brand name and read it back. If it would work unchanged for a competitor, you wrote a category advertisement rather than yours.
The three failures, in the order they arrive
The first two show up in week one. The third arrives for your whole category at once, and it arrives late.
The script is the average
Familiar language, scrolled past
Nothing is at stake
No verifiable person behind the claim
Everyone converges
The look stops reading as modern
Failure two: nobody is standing behind the claim
Advertising works partly because someone is. A named person. A real room. A product handled by hands that have handled it before.
Fully synthetic ads remove all of that at once. No verifiable person, no real environment, no demonstration that could not have been fabricated.
The viewer may never articulate the problem. The credibility signal is missing and the response rate reflects it.
That is why generated presenters do fine on informational scripts and poorly on emotional testimony. Explaining how something works asks nothing of the viewer's trust in the speaker.
Generated product footage fails the same test for a different reason. Models are good at plausible objects and bad at consistent specific ones. Your packaging shifts shade between shots.
Where a generated presenter works, and where it reliably disappoints
The split is not quality. It is whether the script asks the viewer to trust the speaker personally.
| What the script asks of the viewer | How a synthetic presenter handles it | |
|---|---|---|
| Explaining a mechanism | Follow an explanation, trust nobody | Works. The speaker is a narrator |
| Walking through a comparison | Weigh two options on the facts given | Works. Nothing rests on who is talking |
| First-person testimony | Believe this happened to this person | Fails. A face performing a feeling it lacks |
| A demonstration in hands | Accept the object is the real one | Fails. It drifts between shots and is noticed |
| A checkable claim | Trust that somebody stands behind it | Fails. Nobody does, and it reads that way |
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Failure three: within months the whole category looks like one advertiser
When a category adopts the same tools, the outputs converge. The same visual style, the same pacing, the same caption animation, the same three-act structure.
Convergence is expensive in a channel that rewards novelty. Be the fifteenth ad this week to open with the same generated look and you are borrowing somebody else's identity.
It is not even a competitor's identity. It belongs to the tool. Your viewer has learned to recognise it faster than you can change it.
The defence is keeping the parts that cannot converge. Your customers' phrasing. Your founder's face. Your specific numbers. Your unusual angle.
Watch what happens in any category that has been through this. The generated look stops reading as modern and starts reading as advertisement. The whole cohort loses hook rate together.
Being early is an advantage that expires, usually faster than the contract that bought it. Check your own feed for the tell. Open the ad library for three competitors and count how many share your caption animation.
Give it the labour, keep the judgement, and know which is which
None of this is an argument against the tools. It is an argument about which jobs to give them.
They are excellent at labour. Transcribing a take accurately. Assembling a first cut. Timing captions to the word. Finding a clip that matches a phrase. Producing the eighth variant of something you already decided was right.
All of that is work with no taste in it. It is also the work that caps how much creative a small team can test.
They are bad at deciding what to say, what is true about your business, and what your customer is afraid of. They are also bad at knowing when to stop.
An unsupervised pipeline produces twelve competent ads a week and no argument at all.
The two lists, and the line between them
Everything on the left has no taste in it. Everything on the right is why your ad is different from a competitor's.
| Give it to the machine | Keep it in human hands | |
|---|---|---|
| Transcription | Accurate, fast, and correctable | Deciding what was worth saying |
| Assembly | Cuts, timing, caption sync | The opening line and what to cut |
| Footage | Matching a clip to a phrase | Which phrase deserved a picture |
| Variants | The eighth version of a decided idea | Whether the idea was worth eight |
| Specifics | Nothing. It cannot know them | Price, objection, the customer's own words |
Automate the left column. Keep the right one. That is the whole design.
Cutroom takes the labour list and hands back the judgement list untouched. You record the take, so the voice, the room and the specifics are yours from the first frame.
It transcribes, drafts the edit, finds the footage, times the captions and returns a finished 9:16 MP4. Every change is made on the transcript rather than on a timeline.
That is the answer to all three failures on this page. A real person said the words, so failure two never arises. The words are yours, so failure one never arises. The face is yours, so failure three cannot reach you.
A fully generated pipeline cannot make that claim. What it removes is the exact part that was working.
The facts to plan around. Somebody has to be on camera. Three minutes is the upload ceiling. Every export is a 9:16 MP4, and there is no timeline underneath.
The machine never decides the argument. Deciding what is worth saying is still the part of your week that pays.
Questions people ask
- Are generated presenters worth using at all?
- Yes, for informational scripts where the speaker is a narrator rather than a witness. Explainers, comparisons and walkthroughs work. First-person emotional claims from a synthetic face are the case that reliably disappoints.
- Can I fix a generated script by editing it?
- Usually you fix it by replacing most of it. Keep the structure if it is sound, then rewrite the sentences with real detail from your own business. The parts that survive editing tend to be the skeleton, not the language.
- Do platforms penalise AI-generated content?
- The bigger risk is audience response rather than policy, though disclosure requirements exist in some markets and are worth checking. An ad that viewers scroll past gets throttled by performance long before any policy question arises.
- So what should stay human in the process?
- The angle, the opening line, the claim, the proof and the decision about what to cut. Those move results. Transcription, assembly, caption timing and variant production are labour and can safely be automated.
- Who should not buy any of these tools this quarter?
- Anyone whose last ten ads failed for reasons of offer or audience. Faster production multiplies whatever you already have. Fix the offer first, then come back and multiply the better one.
Give the machine the labour and keep the judgement. You record the take, so the voice, the room and the specifics stay yours. A generated pipeline cannot say that.