Guide
Yes, end to end. Two of the four jobs are still yours
A machine takes a recorded take and returns a finished vertical ad today. The assembly is the dependable part of that. Deciding what to say is not automated. Judging whether it worked is not either. This page splits the question into four separate jobs and prices the two that are now mechanical. By the end you will know which part of your week you can hand over this month.

- Separate jobs inside one question
- 4
- That are reliably mechanical today
- 2
- Free credits, 7 day trial, no card
- 300
Assembly is solved, capture is half solved, judgement is not automated at all
Making a video ad is four jobs, not one. Deciding what to say. Capturing something to look at. Assembling it into a finished file. Judging the result.
Assembly is largely solved. Capture is partly solved and unevenly. Writing is assisted rather than automated. Judgement is not automated at all.
Asked of the whole bundle at once, the question collects a yes and a no. It lands on nothing. Asked of each job separately it has four clean answers.
Judgement stays put for a concrete reason. It runs on facts about your business that sit in no training set. Which objection blocks the sale. What a winner has to return. Whether the version in front of you sounds like a person or a brand.
The two mechanical jobs are the ones with a price list, and Cutroom takes them whole. One recorded take goes in. A finished 9:16 file comes out.
Four jobs, four different answers
Two you can hand over today. The first and the last decide whether the ad works, and both live in your inbox.
Decide what to say
Yours. Nobody else has the facts.
Capture
Partly. Fails on your product.
Assemble
Solved. Transcribe, cut, caption.
Judge the result
Yours. Your money, your bar.
The chores that used to be billed by the hour now have a price list
These are the parts that were invoiced as craft. None of them is a creative act. Every one was a chore somebody dreaded.
Notice the shape of the list below. Each item transforms material you already supplied. That is why the failure rate is low and the prices are countable rather than negotiated.
A batch is 100 credits whatever the take contains. An export is 20 credits per output minute. So a thirty second ad is 110 credits start to finish, against the 2,500 credits Basic carries each month.
- Speech to text with timing accurate to the individual word, from a take recorded on a phone.
- Cutting a three minute ramble down by removing pauses, filler and the two false starts before you found the sentence.
- Timing and styling captions across six shipped packs, including emphasis on the one word carrying the claim.
- Finding stock footage that matches a phrase, out of libraries nobody has an afternoon to browse.
- Producing a 9:16 MP4 at the right dimensions and loudness, at 20 credits per output minute.
- Generating a music bed from a written description, at 20 credits a track.
- Turning a photo and a voice track into a lip-synced presenter, at 320 credits per rendered minute.
What the mechanical steps cost, in credits
A thirty second ad is one batch and half an export minute: 110 credits. Basic carries 2,500 a month.
- One music track20 CR
- Export, per output minute20 CR
- Voice-over, per minute70 CR
- One batch, whole take100 CR
- Avatar, per rendered minute320 CR
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Your packaging changes shade between shots, and a customer notices
Generated footage of your actual product is the clearest half-solved case. Models are good at plausible objects. They are bad at consistent specific ones.
The packaging shifts colour between shots. The logo becomes a smudge that is nearly your logo. For anything a customer will recognise, this stays unreliable. The failure is not subtle.
Synthetic performance is the second. A generated presenter reading an informational script is fine. The same presenter delivering first-person testimony usually is not.
Viewers are unusually good at spotting a face performing a feeling it does not have. That failure is social rather than visual. No amount of resolution repairs it.
Scripts are the third. A language model writes competent, well-structured, entirely generic ad copy in seconds. What it cannot supply is why the second version changed, what the return rate is, or the sentence a customer used last Tuesday.
The three share a shape. Each passes a casual look. Each fails a paying audience. They break exactly where a viewer needed to trust something specific.
Nobody can automate knowing which objection is blocking the sale
That knowledge comes from talking to customers, reading cancellations and losing deals. It is the highest-value input any ad has. It exists in your inbox and your head, nowhere else.
The offer is the second thing no editing tool touches. What you sell, at what price, with what guarantee. The offer outperforms the execution more often than anyone in advertising admits out loud.
The verdict is the third. What counts as a winner, at what volume, against what threshold. That is a business decision with your own money on the other side of it.
Taste is the fourth, in the narrow sense that matters here. Knowing which of two openings sounds like a person and which sounds like a brand.
Models trained on the average of the internet drift towards the average of the internet. That register is exactly what people scroll past.
Hand over the two mechanical jobs, keep the two that decide the result
Cutroom is built to exactly that split. You supply the take, the claim and the verdict. It supplies the transcription, the cut, the captions, the footage search and the export.
One talking-head take of up to three minutes goes up. A finished 9:16 MP4 comes back. Nothing in the middle lands on your evening.
You direct it on the transcript rather than a timeline. Highlight a phrase and a clip lands over exactly those words. Delete a line and the cut rebuilds around the gap. Change the pace and the machine re-cuts the whole thing.
A batch is 100 credits. An export is 20 credits per output minute. A thirty second ad is 110 credits, against the 2,500 Basic carries each month.
The gain is more attempts at the same price. Published benchmarks put the winner share at 5 to 8 percent, per Motion's analysis of 550,000+ Meta ads. Six attempts a month is not a creative problem. It is an under-sampling problem.
Three facts a buyer needs going in. Three minutes of source per take. 9:16 output. Somebody willing to speak to a camera in a real room. Those are the shape of the format.
Fourteen attempts, one expected winner
At the middle of the published 5 to 8 percent range. The cost of attempt fifteen decides more than the polish on attempt one.
Questions people ask
- Do AI-made ads perform worse than human-made ones?
- There is no reliable evidence either way, because the variable that decides performance is the message rather than the production method. What does perform badly is generic content, and automated pipelines drift towards generic unless you feed them something specific. Judge the output, not the process.
- Will the platforms penalise AI-generated video?
- Policies differ and they change, so check the current rules for each placement you run on. The broad direction is towards labelling requirements for synthetic media, particularly realistic depictions of people, rather than outright suppression. Assume you will need to disclose.
- Can I use a generated presenter instead of filming myself?
- For informational scripts, often yes. For first-person testimony about how something made you feel, it usually reads false and costs the credibility the format depends on. A workable rule is a synthetic presenter for explanation and a real person for claims.
- What is the minimum a human still has to do?
- Decide what to say and decide whether it worked. Everything between those two can be handed off. Neither of them can. Both need facts about your customers that exist nowhere in a model's training data.
- Is it right for every team?
- It is built for teams where somebody will speak to a camera. If nobody will, book a creator per deliverable and compare the options on turnaround and revision rounds.
One recorded take goes in and a file you can upload comes out, with the cut, the footage and the captions already made. Most tools hand back a step.