Guide
Creative testing is running ads to find the one that works
Creative testing means running several versions of an ad against each other so the account tells you which one works. The winner gets the budget. The rest get switched off. This page covers the definition in plain words, one test written out with real numbers, and the arithmetic that decides whether your testing is worth doing at all. By the end you will know how many ads your month has to produce.
- Creatives that become winners
- 5 to 8%
- Expected winners from four ads a month
- 0.26
- Credits for twelve ads from six takes
- 720
The definition: several ads, one question, and the account decides
You make more than one ad. You put them live at the same time, to the same audience, with the same offer and the same budget. You wait until each has had enough impressions to mean something. Then you keep the one that performed and turn the others off.
That is creative testing. Everything else written about it is a variation on those five steps.
Two conditions make it a test rather than a guess. One variable changed between the ads. A verdict rule written down before anything went live.
People know about the first. The second gets skipped, and skipping it is how a test turns into a Friday argument about whether 1.8 percent beats 1.6 percent.
Write the rule first. Something like this. After 8,000 impressions each, the ad with the lowest cost per click wins. Inside ten percent of each other, the test is a tie and neither gets promoted.
A rule written before you have seen the numbers is a decision. A rule written after is a justification.
One test, written out, with the numbers
A magnesium supplement. The question is which opening line earns the third second. So the only thing that changes between the three ads is the first sentence.
Ad A opens with the problem. He says he was waking up at three every night for a month. Ad B opens with the product. He holds the tub and names it. Ad C opens with a number. He had tried four other tubs and this was the fifth.
Everything after the first four seconds is identical footage. Same body, same cutaways, same captions, same offer, same length. That is what makes it a test.
The rule, written on the Monday before launch. Run to 10,000 impressions each at the same budget. Verdict on hook rate. A difference under fifteen percent counts as no result.
Thursday. A has a hook rate of 31 percent, B has 19 percent, C has 29 percent. A and C are inside fifteen percent of each other, so the finding is not that A won. The finding is that opening on the product loses, and it lost by a lot.
That is a real result and it is worth more than a winner. It applies to every ad you make from now on, rather than to one creative that fatigues in three weeks.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Roughly 5 to 8 percent of creatives become winners, and that changes the maths
Published benchmarks put the share of ad creatives that become winners at roughly 5 to 8 percent, from Motion's analysis of 550,000+ Meta ads. Somewhere between one in thirteen and one in twenty.
Take the middle of that range and do the arithmetic on your own output. Four ads a month at 6.5 percent is 0.26 expected winners. You are running a process designed to find something you gave yourself a one in four chance of holding.
Twelve ads a month is 0.78. Twenty is 1.3.
None of that is a comment on how good your ads are. It is a sampling problem. Sampling problems are fixed by sample size, not by trying harder.
This is why creative testing feels broken for small teams. It works exactly as well for them. They are drawing four tiles instead of twenty and concluding the bag is empty.
Read those numbers the other way and they are a relief. Four ads a month producing nothing is the expected outcome, not evidence about your judgement.
Expected winners a month, by how many ads you run
At the middle of the published 5 to 8 percent range. Nothing here is about ad quality. It is how many draws you took.
- Four ads a month0.26 winners
- Twelve ads a month0.78 winners
- Twenty ads a month1.3 winners
The bottleneck is not the analysis. It is how many ads exist.
Most advice about creative testing is about the reading. Which metric, how many impressions, when to call it. That advice is fine. It is not where the money is being lost.
The money is lost upstream, where somebody decides which ideas get made. Four slots and eleven ideas means seven ideas were killed by a booking calendar rather than by an audience.
The seven that got killed were not the safe ones. Nobody spends their one production slot on the strange angle. They spend it on the version most likely to be acceptable, which is also the version most likely to be average.
So the real cost of expensive production is not the invoice. Testing becomes rationing, and rationing selects against exactly the ideas that produce outliers.
The fix is arithmetic rather than inspiration. Make the next ad cheap enough that the strange angle gets made anyway. The sample size problem and the selection problem go away together.
That is the whole argument for changing how ads get produced. Not that machine assembly makes better ads. That it makes the eleventh one exist.
Twenty creatives, one winner
At the cautious end of the published range. Every tile you never made is a draw you never took, and the strange angles are the ones that get cut first.
Twelve ads a month from six takes is 720 credits
Here is what the volume costs when the assembly is not done by hand.
Record six continuous takes on a phone, up to three minutes each, in one morning. That is the only part that needs a person and a room.
Each take gets a batch at 100 credits. That covers the transcript, the director pass and the b-roll search. The cut comes back made. Trimmed, covered, captioned and paced.
You direct it on the transcript. Highlight a phrase and a clip lands over exactly those words. Delete a line and the cut rebuilds. Change the pace and it re-cuts.
That is where the second and third ad from one take come from. A new opening plus a different set of highlights is a genuinely different creative.
Twelve thirty second exports at 20 credits per output minute is 120. Six batches is 600. Total 720 credits, against the 2,500 Basic carries at $39.99 a month.
The trial is 7 days and 300 credits with no card. That is two takes and four finished thirty second ads, enough to run one real three-way hook test before any money moves.
Questions people ask
- How many impressions before I call a result?
- Enough that the difference you are reading is larger than the noise. For hook rate on a cheap metric, several thousand impressions per creative is usually workable. Write the number down before launch, because the number you pick afterwards will be the one that agrees with you.
- Should I test one variable or a whole new concept?
- Both, in different tests. One variable tells you why something worked and gives you a rule you can reuse. A whole new concept tells you whether a different angle exists at all. Concept tests find outliers, variable tests explain them.
- How is this different from A/B testing a landing page?
- The mechanics are the same and the win rate is not. A page test moves a conversion rate a few percent. A creative test is looking for a small number of outliers among many failures, which is why volume matters so much more here than it does on a page.
- When does a winning creative stop winning?
- Sooner than you plan for. Frequency climbs, the audience has seen it, and performance decays. That is why the pipeline matters more than the winner: you need the next test already running when the current one fades.
- Is volume right for every account?
- It pays once each ad can be given enough impressions to answer. If your spend is small enough that twelve creatives never reach a readable sample, run fewer ads for longer and test offers and audiences instead.
Write the verdict rule before launch, then count how many ads exist. Cutroom makes the eleventh one cost 110 credits, which is the number that moves the count.