Feature
AI spokesperson for ecommerce: one face, forty listings
A catalogue is a scheduling problem before it is a creative one. Forty SKUs means forty scripts. Nobody stands in front of a camera forty times. A rendered spokesperson removes that constraint at 320 credits a finished minute. Here is what one SKU ad costs, how a single render fronts many listings, and where to spend the credits first.

- Per minute of spokesperson render
- 320 CR
- A forty second clip with a generated read
- 260 CR
- Credits a month on Basic at $39.99
- 2,500
What one SKU costs, before you multiply it
One forty second spokesperson clip with a generated read is about 213 credits of render plus 47 of voice-over.
Call it 260. Batching that clip into a real ad with captions and product footage adds 100. The export adds about 14.
So one finished spokesperson ad for one SKU is around 374 credits.
Basic is $39.99 a month for 2,500 credits. That is six of those. Premium is $79.99 for 5,000, which is thirteen.
A top-up is $15 for 1,000 credits. That buys two and a half more.
The trial is 300 credits with no card. Enough to take one SKU all the way to a finished file before you decide.
Filmed source takes run on the same plan at 100 credits a batch. The same budget stretches further where a camera is possible.
Both numbers matter. No catalogue is written in a single register.
Finished ads per month, same budget
Same spend, two production routes, one account. The mix is the decision.
- Filmed source, Basic plan21 ads
- Rendered spokesperson, Basic plan6 ads
One render can front many SKUs, because the footage underneath changes
There are no hands in a rendered frame. No unboxing, no label turned to the lens, no pump on the back of a wrist.
That sounds like a wall. It is actually what makes the render reusable.
Shoot one session of product footage on a phone. An afternoon covers a catalogue.
Then render the spokesperson once for a script shape you use often. An offer. A spec comparison. A restock.
In the editor you highlight the phrase that names the product. A clip lands over exactly those words.
Swap that clip in any spot, or search millions of free ones. The same rendered read then fronts three SKUs with different footage under each.
Change the caption headline and the pack. The two ads read as siblings rather than as copies.
That is the version of this that pays back. Rendering a fresh face per SKU and posting it bare is the version that does not.
None of it needs a second tool. The swap is a highlight in the same editor that made the cut.
One render, many SKUs
The expensive part is the render. The cheap part is swapping what sits under the words.
Shoot product footage once
Phone, one session, all SKUs
Render the spokesperson
320 CR a minute
Batch it
100 CR, transcript and cut
Swap the clips per SKU
Highlight the phrase, drop the shot
Export each
20 CR per output minute
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Consistency is the reason ecommerce teams want this
The same presenter fronts January and June. Nobody is rebooked, restyled or unavailable.
That matters more for a catalogue than for a single product. A recognisable face across forty listings does work no individual ad does.
It also survives staff turnover. A filmed spokesperson who leaves takes every ad they appear in with them.
There is a quieter benefit. A fixed presenter means the only variable between two ads is the script.
So your test results tell you which script won. With filmed creators, every ad varies in lighting, energy and room.
The winner you find there may be a person rather than a message. That is an expensive thing to learn late.
The permission rule keeps this safe. The portrait is yours, or belongs to somebody who gave explicit written permission. That permission should name how long it runs.
A stock or image-model presenter sidesteps the question, at the cost of a face nobody recognises.
Decide it once at the start. Changing presenter across a catalogue costs more than choosing carefully did.
Which sentences to render and which to film
There is a rule teams land on after a month of running both. It fits in one line.
If the sentence contains a number or a policy, render it. If it contains an adjective about how something feels, film it.
Render: free shipping over forty euros, ships Tuesday, two hundred servings a tub, restocked this week.
Film: how heavy it is, how the fabric moves, how it smells when the lid comes off, the moment somebody sees the result.
Scale needs a hand next to it. Held next to a hand is a fact. Described as compact is a claim.
First-person testimony about a physical experience belongs on a phone too. A rendered face performing a memory it never had reads wrong.
Long scripts belong to neither. At 320 credits a minute, ninety seconds is 480 credits of render and loses viewers anyway.
Forty seconds is the working ceiling for a rendered read, which is also the length a feed rewards.
That split decides where the 320 credits a minute earns its price, and it takes ten minutes with a pen.
Where to spend the credits first
Buy the render first for the scripts nobody will film. Policy updates, shipping cutoffs, spec comparisons, restock announcements.
Those are the ads that never get made otherwise. They are the ones the render makes cheap.
If you already pay creators for footage, you own faces and hands. Batch those takes at 100 credits each and spend the difference on more attempts.
Published benchmarks put the winner rate at roughly 5 to 8 percent, from Motion's analysis of 550,000+ Meta ads.
Cost per attempt is the number that decides whether you reach a winner. The mix matters more than the tool choice.
The plain facts to plan around. Output is 9:16. A filmed source take caps at three minutes. Rendered audio caps at five.
There is no timeline, by design. Every lever is a decision about the ad.
What you get on both routes is a finished vertical ad rather than a clip to take somewhere else.
For a catalogue that is the part that decides whether forty listings ever get forty ads.
Questions people ask
- Can the spokesperson show the product?
- Not in the render. Film the product separately on a phone and place that footage under the phrases where the spokesperson describes it. One product shoot serves every script you render afterwards.
- How many SKUs can one presenter cover?
- As many as you like, because the presenter is not tied to a product. The limit is credits: at roughly 374 credits per finished ad, Basic covers six a month and Premium thirteen.
- Is a stock presenter better than using my own face?
- It is safer and less memorable. A stock or image-model face removes the permission question and cannot leave the company. Your own face carries more trust if customers already know it.
- Who should not buy this?
- Anybody whose catalogue sells purely on texture and scale. Put that budget into filmed takes at 100 credits a batch and keep the render for announcements and specs.
One face across a catalogue, different footage under every script, and a finished 9:16 file at the end of each one.