Guide

One variable, enough volume, and a rule written before launch

Change one variable at a time. Give each version enough impressions to mean something. Write the verdict rule down before anything goes live. Almost every failed creative test breaks one of those three, usually the third. This page covers what to test first, how much volume a result needs, and why your cadence is capped by supply. By the end you will know how many creatives your month has to produce.

A person working on a laptop at a desk with snacks, emphasizing productivity and technology.
Photo by Kampus Production on Pexels
Rules, and the third is the one that breaks
3
Tested creatives per expected winner
14
Published winner share across Meta ads
5-8%

Two ads that differ in four ways teach you nothing about any of them

A test is a comparison, and a comparison needs one difference. Two ads with a different hook, a different pace, a different caption style and a different ending produce a result you cannot act on.

You will know one performed better. You will not know which change did it. So the next round starts from guesswork again.

The variable worth testing first is almost always the opening. It has the most variance in outcome. It also has the lowest cost to change.

Record six openings against the same body in one sitting and you have six creatives from one recording session. Six different shapes, not six rewordings of one. Near-identical variants produce near-identical numbers.

The second variable worth testing is the offer, and it is not a creative test at all. It is the largest lever in the account and it belongs in a separate experiment.

The order to test in, and why

Each step down costs more to change and moves the number less. Most accounts start at the bottom and never reach the top.

  1. The offer

    Largest effect, not a creative test

  2. The opening

    Most variance, cheapest to vary

  3. The claim and proof

    Needs new material

  4. Pace and captions

    Real, and smaller

Underpowered tests are how an account learns what is not true

A creative that wins on four hundred impressions has told you nothing. Noise at that volume is larger than any effect you are looking for. Acting on it is worse than not testing.

Decide in advance how much each arm gets and let it run. The most common failure is killing an arm on day one, because it looked bad before lunch.

Judge on the fastest reliable metric available rather than on conversions. Conversions arrive too late and too sparsely for most small accounts to read.

Hook rate reads the opening. Hold to the halfway point reads the body. Cost per landing page view reads intent. Conversions read the offer, eventually.

That ladder lets you kill a bad opening in two days. The alternative is three weeks waiting for a conversion signal that never becomes significant.

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

Write the verdict rule before launch, or you will find one afterwards

This is the rule that breaks most often, and it breaks silently. Without a rule written down, the winner gets decided by whoever opens the dashboard first and whichever number flatters their preference.

The rule has three parts. The metric. The threshold. The volume at which you will read it. One sentence, written before anything goes live.

Something like this. At 8,000 impressions per arm, the higher hold rate to fifty percent wins. Inside a tenth of each other, the test is a draw and we keep the cheaper one to produce.

A draw is a legitimate result and most tests produce one. Recording draws is how you stop retesting the same idea every quarter.

Log the result somewhere permanent, in one line. What changed, what won, by how much, at what volume. Six months of that log is worth more than any benchmark article.

  • The metric, named before launch: hook rate, hold to halfway, cost per landing page view.
  • The threshold: the margin below which you call it a draw rather than a win.
  • The volume: impressions per arm before anybody is allowed to look.
  • The consequence: what actually happens to the loser, and who does it.

Your test cadence is capped by supply, not by ambition

Published benchmarks put the winner share at 5 to 8 percent, per Motion's analysis of 550,000+ Meta ads. At the middle of that range you need roughly fourteen tested creatives to expect one winner.

That makes the binding constraint obvious. Six creatives a month is not a creative problem. It is an under-sampling problem, and no amount of care on any single ad fixes it.

So the question is what your production route yields per month. Then whether that number reaches fourteen inside the window where you need a new winner.

Compare routes on turnaround, revision rounds, cost per edit and consistency rather than on quality claims. Those four decide the count.

Assembly is the line with a fixed rate available. One take of up to three minutes returns a finished 9:16 MP4, directed on the transcript rather than a timeline. Highlight a phrase and a clip lands over exactly those words. Delete a line and the cut rebuilds. Change the pace and it re-cuts.

A batch is 100 credits. An export is 20 credits per output minute. A thirty second ad is 110 credits, against the 2,500 Basic carries monthly.

Fourteen tested creatives, one expected winner

At the middle of the published 5 to 8 percent range. Count what your current route yields per month and compare it to this grid.

Carries the spendPaid for in full

Make the next attempt cheap and the count moves, which is the only lever that reaches results

Every route to volume has a cost. Recording yourself costs willingness. Commissioning creators costs money and days. Customer content costs predictability.

Assembly is the one line anybody will quote a fixed rate for. A thirty second ad is 110 credits here, whatever the take contains.

Basic carries 2,500 credits a month at $39.99. That is roughly twenty-two finished ads, more attempts than most teams manage in a quarter. Premium is $79.99 for 5,000.

The trial is 7 days and 300 credits with no card. Enough for two takes and a handful of thirty second exports before any money moves.

Three facts to build the plan around. Three minutes of source per take. 9:16 output. Six caption packs, with colour, weight, size and position adjustable inside each.

A test needing four shapes, a second location or an unfilmed demonstration goes to a videographer. Run the cheap route for the volume around it, because the volume is where the winner turns up.

Questions people ask

How many creatives should I test at once?
As many as you can give meaningful volume to, and no more. Splitting a small budget across eight arms produces eight underpowered results. Two or three well-funded arms teach you something. Eight starved ones teach you noise.
Should I test in one ad set or several?
One ad set, same audience, same budget, same offer, for a clean comparison. Testing across ad sets introduces audience differences you cannot separate from creative differences, which puts you back to guessing which change did it.
How long should a creative test run?
Until each arm has the impressions you named before launch, and not a day less. Time is the wrong unit, because a week on a small budget can be a hundredth of the volume a week on a large one delivers.
What do I do with the losers?
Turn them off and write down why, in one line. The log is the compounding asset here. Without it, teams retest the same idea two quarters later and pay again for an answer they already bought.
Is testing right for every budget?
It pays once two arms can each get meaningful volume. Below that, put everything behind one creative and spend the effort on the offer. Start testing when the budget can pay for a real answer.

One variable, enough impressions, a rule written first. Then make the eleventh attempt cost 110 credits instead of an evening, which is the part most tools still leave with you.

Start with one take300 free credits · no card · cancel anytime