Guide
Five rules, and the fifth decides the other four
Most creative testing frameworks are four rules long. A falsifiable sentence per creative. Verdict conditions written before launch. Kill on cheap metrics, promote on expensive ones. Log every row, including the ties. Each of those is here with its thresholds written out. Then comes the fifth rule, which decides whether the other four were worth writing down.

- Expected winners from four creatives a month at a 7% rate
- 0.28
- Creatives in a hundred that are supposed to fail
- 93
- Rules, and the fifth one is the only operational one
- 5
Rule one: if you cannot write the sentence, the creative is decoration
Before a creative goes live, write the claim it tests. Not the description, the claim.
Try 'leading with the price objection beats leading with the outcome'. Or 'a demonstration in the first two seconds holds more viewers than a face'. Both are falsifiable, which is the point.
If you cannot write that sentence, the creative teaches you nothing however it performs. This rule alone removes a third of most testing calendars.
Then write the layer next to it. Hook, angle, format, proof, length, presenter. Six creatives in the hook layer are one experiment. Six across six layers are a survey. A survey is what you want in month one.
The five rules, in the order they fail
Teams adopt rules one to four and skip five. Then the framework runs on four creatives a month and decides almost nothing.
Write the sentence
Falsifiable, or it is decoration
Set the verdict conditions
Metric, threshold, volume floor
Kill fast, promote slowly
Cheap metrics out, expensive metrics in
Log every row
Including the ties nobody records
Feed it enough attempts
The rule everybody skips
Rule two: the volume floor is the clause everyone leaves out
Write down three items before launch. The metric, the threshold, and the volume required to call it.
A conversion rate estimated from ten conversions has a confidence interval you could drive a bus through. Skip this step and you are arguing about noise by Thursday.
A common working rule is fifty to a hundred conversions per arm before treating a difference as real. If that is unaffordable, judge on an earlier and cheaper event instead.
Three-second view rate, hold rate to the halfway point and click-through rate accumulate hundreds of times faster than purchases. Use them to eliminate obvious losers. Reserve conversion-level judgement for survivors.
Write the early threshold as a number. Bottom quartile of the batch on three-second view rate after fifty thousand impressions is a rule. Looks weak is not.
One more clause is worth writing down. Who is allowed to call it. A verdict announced by whoever opened the dashboard first is how a threshold gets renegotiated on a Friday afternoon.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Rule three: kill fast and promote slowly, because the asymmetry is enormous
Killing a bad creative early costs a small amount of information. Promoting a false winner early costs a budget cycle. It also costs a wrong lesson repeated across the next quarter.
So kill on the cheap early metrics and promote only on the expensive late ones. Never the other way round.
A creative whose three-second view rate sits far below the batch median after a meaningful number of impressions can go now. One whose cost per acquisition looks brilliant after six conversions stays where it is.
Remember what the base rate implies. Published benchmarks put the winner share at 5 to 8 percent, per Motion's analysis of 550,000+ Meta ads.
Most of what you launch is supposed to fail. A testing process that feels brutal is usually a testing process that is calibrated.
Seven in a hundred, at the optimistic end of the range
Ninety-three of these are meant to die. A process that kills fast is not harsh, it is calibrated.
Rule four: log the ties, or the same dead layer gets tested twice a year
One row per creative. The hypothesis, the launch date, the variant details, the result, the verdict, and one sentence on what it taught you. A spreadsheet is enough.
Without it you will retest the same hook three times a year and never notice. Every new hire starts from zero.
The log is also where you compute your own win rate. After forty or fifty rows you can stop planning against a published benchmark and plan against yourself.
Keep the ties in it. A test where two variants landed on top of each other feels like a failure. It is a finding. That layer does not move your audience.
Most teams delete those rows. That is how the same dead layer gets tested twice a year for three years, by three people who each thought it was a fresh idea.
Rule five: four creatives a month is 0.28 expected winners, process or no process
Expected winners equals your win rate times your monthly output. At 7 percent, four creatives is 0.28 expected winners.
A disciplined process running on four a month will make excellent decisions about almost nothing. So the last rule is operational rather than analytical.
Get the cost of an attempt low enough that the framework has something to chew on. That is what Cutroom is for. One take in, a finished 9:16 MP4 out, directed on the transcript rather than assembled on a timeline.
Six hypotheses become six exports from one recording session. Trim the opening. Delete a line. Change the pace. Each variant costs 20 credits per output minute, against a 2,500-credit month on Basic.
Nothing else takes you from a phone recording to twelve testable files in an afternoon. An editor prices each variant as a job. A timeline charges you an evening each.
The facts to plan around. The source is a person speaking, up to three minutes. Every export is a 9:16 MP4. Every hypothesis about an opening, a length or a pace is one export away.
Expected winners a month, at a 7 percent rate
Multiply your output by your rate. A rigorous process on the top bar is rigour applied to nothing.
- 4 creatives0.28
- 12 creatives0.84
- 30 creatives2.1
Questions people ask
- Should I run a formal A/B test or launch everything?
- For creative, structured launching with pre-committed decision rules is usually more practical than strict controlled testing, because the delivery system is not a neutral splitter. What matters is that the threshold and the volume floor were set before you saw any data.
- How many variables can one creative test?
- One, if you want a transferable lesson. You can launch creatives that differ in several ways while exploring, but be honest that you are hunting for a winner rather than learning a rule, and do not write the lesson down.
- What if a creative wins but I do not know why?
- Run the deconstruction as a test. Rebuild it with the suspected ingredient removed. If performance holds, that was not the ingredient. This is slower than guessing and it is the only way a win becomes a repeatable rule.
- How long should a test run?
- Until it hits the pre-committed event count, or until an early metric has clearly eliminated it. Calendar-based test lengths are borrowed from campaign reporting and they cause both premature verdicts and pointless waiting.
- Who should not adopt this framework at all?
- Anyone spending too little to reach a volume floor on any metric. Below that point you are running a diary rather than a test. Keep rules four and five, and let the rest wait until the spend can pay for a verdict.
Rules one to four cost nothing to adopt. Rule five costs one take a week, because the finished vertical ad comes back in minutes instead of a booking.