Feature
AI UGC avatar ads: the face says it, the footage shows it
UGC converts because somebody shows you something. A rendered presenter says anything and holds nothing. That is a shape to build around, and the build takes one step. Here is what the render gives you, the hybrid that puts real footage under the words, and what twenty finished attempts cost.

- Hands available in a generated frame
- 0
- Per minute of presenter render
- 320 CR
- Runtime the editor can cover with footage at normal pace
- 45%
What renders, and the one thing that does not
You give it a front-on portrait and a voice-over. Back comes a 9:16 MP4 of that person speaking, framed head and shoulders.
The mouth moves with the audio. The eyes hold the lens. The room behind stays exactly as it was in the still.
There are no hands. The presenter cannot unbox anything, tilt a label toward the camera, or pump a sample onto a wrist.
Most UGC scripts that convert contain one of those moments, usually inside the first four seconds.
So read your script once. Mark every sentence that needs something shown.
Those marked sentences are the footage list. Everything else is the render list.
That single split is the whole method. It takes about ten minutes with a pen.
Nothing about render quality changes that split. It is how you get a UGC ad out of a format built for words.
- Available: a face, eye contact, a fixed room, a voice in sync
- Not available: hands, a product, a second angle, movement through a space
- Cost: 320 credits per finished minute, plus 70 a minute if the read is generated
- Trial: 300 credits, roughly one minute of render, no card
The hybrid: rendered voice on top, real footage underneath
The version that converts uses the presenter for the words and real footage for the showing.
Hand the rendered clip to the editor as a source take. One batch is 100 credits.
It comes back as a finished ad you direct on the transcript.
Highlight the phrase this is the applicator. A clip lands over exactly those words. The face steps out for that beat and comes back after it.
At normal pace the compiler covers up to 45 percent of the runtime. At fast pace it goes to 52.
So around half the ad is footage rather than portrait. That is the difference between a rendered ad and a rendered stare.
Swap the clip in any spot. Search millions of free ones. Or upload the twenty seconds you shot of your own product on a kitchen counter.
Captions come from six style packs, burned in. Most of the feed is watching on mute.
Every other presenter tool hands this step back to you as homework.
How much of the finished ad is not the face
At normal pace the compiler covers up to 45 percent of the runtime with footage. The portrait carries the rest.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
The cost arithmetic, and where to point the render
Testing creative is a volume problem. Published benchmarks put the winner rate at roughly 5 to 8 percent, from Motion's analysis of 550,000+ Meta ads.
At that rate you need somewhere between thirteen and twenty attempts to expect one winner.
Twenty forty-second rendered clips with generated reads is about 5,200 credits.
Twenty filmed takes batched and exported is about 2,280 credits. Basic at $39.99 covers most of that.
Point the render at the scripts a camera cannot cover. Announcements, offers, spec walkthroughs, anything nobody will read out loud.
Those are the ads that never get made otherwise. The render makes them cheap and repeatable.
Film the testimony and the demonstrations. Batch them at 100 credits each. Put those clips under the rendered lines.
Both halves live in one account and finish in the same export.
That mix is the pattern teams settle on once they have run the numbers themselves.
Twenty attempts, both routes
Twenty forty-second ads. Spend the render credits on the scripts a camera cannot cover.
- Filmed, batched, exported2,280 CR
- Generated presenter with generated read5,200 CR
The problems you will hit before the second render
A script that says look at this. There is nothing to look at. The pause where the demo should be is audible.
A voice-over recorded in a room with hard surfaces. Lip movement follows the waveform. Reverb becomes mushy mouth shapes.
A portrait with a wide smile or a hard shadow across one cheek. Both bake in, then have to move.
First-person testimony. A generated face saying it changed my mornings is performing a feeling it never had. Viewers detect that fast.
And length. Forty seconds is the working ceiling. Past it an even delivery costs more attention than the script earns.
None of these throw an error. They return a clip that feels off for a reason nobody on your team can name.
The cheapest way to find out which one is hurting you is to render twenty seconds rather than ninety.
Watch it on a phone with the sound off. If it survives that, render the rest.
If it does not, the fault is in the script almost every time.
Buy the product footage once, then reuse it under every script
If the ad lives or dies on the unboxing, film thirty seconds of hands and product. A phone on a stand does it.
That footage is reusable across every ad you make afterwards. It is the first purchase rather than the last.
Then render the presenter for the scripts that are pure information. Drop the bought footage under the phrases that need it.
Highlight the phrase, the clip lands, the ad is finished. No timeline. No round trip through another tool.
The plain facts to plan around. Output is 9:16. There is no timeline, by design.
The permission rule has no exceptions. The portrait is yours, or belongs to somebody who gave explicit written permission.
A rendered presenter is a courier for words. Anything that has to be shown needs a camera, and that camera is the phone in your pocket.
Cutroom is where those two halves meet in one file. A render on its own was never going to do that.
Questions people ask
- Can a generated presenter do an unboxing?
- No. There are no hands in the frame. Film the unboxing on a phone, then place it under the phrases where the presenter describes it. One shoot serves every script afterwards.
- Do generated UGC ads underperform filmed ones?
- On explanatory scripts they hold up, because the information is doing the persuading. On testimony a filmed take still wins. Sort your scripts into those two piles and the render stops disappointing you.
- How much of the finished ad shows the avatar?
- Around half, depending on pace. Coverage caps at 40 percent on chill, 45 on normal and 52 on fast. At the fastest setting the presenter holds under half the runtime.
- Who should not buy this?
- Anybody selling something tactile with nobody willing to film anything at all. Spend the first budget on one creator shoot. Then bring that footage back here as b-roll and render the words over it.
The face carries the words. The footage carries the proof. Cutroom is where they arrive as one 9:16 file.