Skip to content

Feature

Film it yourself: the six lines an AI avatar cannot deliver

If somebody is willing to go on camera, filming usually beats the avatar. One phone take of up to three minutes batches for 100 credits and comes back cut, with captions and b-roll on it, and the export adds 20 credits a minute. Here is what the filmed route costs, the six sentences a rendered face cannot carry, and when to render instead.

Close-up of a smartphone displaying time, captured indoors with a blurred background.
Photo by Amar Preciado on Pexels
A filmed forty second ad, start to finish
114 CR
Longest source take the uploader accepts
3 min
Filmed ads a month on Basic at $39.99
21

What filming costs here, so the comparison is honest

Upload one talking-head take of up to three minutes.

One batch of 100 credits covers the transcription, the director pass and the b-roll search.

Export is 20 credits per output minute, so a forty second ad is about 14 more.

That is 114 credits for a finished 9:16 MP4, against roughly 393 for the rendered version of the same script.

Basic is $39.99 a month for 2,500 credits. That is about twenty-one filmed ads or six rendered ones.

The take does not have to be good. You cut three minutes down to forty seconds and direct on the transcript.

Delete a line you do not want and the cut rebuilds around it. That single behaviour is what makes a rambling take usable.

Most editors would have you find that line on a timeline. Here you find it in the words. That is why the take can be rough.

Ads per month on Basic, same $39.99

2,500 credits a month. What you get depends on which route the script needs.

  • Filmed on a phone21 ads
  • Rendered presenter6 ads

Six sentences a generated face cannot carry

First-person testimony. I have used this every morning since March means nothing from a face that has had no mornings.

Anything sensory. How it smells, how heavy it is, how the fabric behaves when you move.

Anything that needs hands. Opening, pouring, holding next to a face for scale, pointing at the small print.

A surprise or a genuine reaction. The half laugh, the pause, the restarted sentence.

Those artefacts are the proof somebody meant it. They cannot be written into a read.

A recommendation with your name attached. If your audience knows you, a render of you saying it is worse than a stranger saying it.

And an apology or a correction. A generated face delivering an apology reads as a company avoiding one.

None of the six are edge cases. Between them they cover most of what a small brand has to say in its first year.

Every one is cheap to film and finishes in the same editor for 100 credits.

  • Testimony, sensory claims, hands, spontaneity, personal recommendation, apology
  • All six are cheap to film and impossible to render
  • None of them improve with a better portrait or a better model

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

The uncomfortable-on-camera objection has a cheaper answer

Most people who say they hate being on camera mean they hate watching themselves do a perfect take.

Nothing here asks for a perfect take. You talk for three minutes, badly, then delete the lines you do not want.

The cut rebuilds around each deletion. There is no gap to patch and no timeline to nudge.

If the face is the objection rather than the performance, point the camera at your hands and the product. Talk over it.

The transcript still works. It comes from your speech rather than from your face.

That gives you a faceless ad for the same 100 credits to cut, plus the export, with your voice carrying the words.

Captions from six style packs carry them again on screen. That is how most of the feed will read them.

The rendered route answers a narrower problem. It is for when nobody will speak at all.

Both answers sit in the same account. Nobody commits to a workflow before they know which one they need.

The three minute take, and what happens to it

Nothing here requires the take to be good.

  1. Talk for 3 minutes

    Phone by a window

  2. Batch

    100 CR, transcript and cut

  3. Delete the bad lines

    The cut rebuilds around them

  4. Cover the rest

    Highlight a phrase, drop a clip

  5. Export

    20 CR per output minute

Volume is the argument that settles it

Published benchmarks put the winner rate at roughly 5 to 8 percent, from Motion's analysis of 550,000+ Meta ads.

That is somewhere between one in thirteen and one in twenty attempts before you expect a winner.

Twenty filmed attempts is about 2,280 credits and fits inside a single month of Basic at $39.99.

Twenty rendered attempts is over 7,800 credits, which is more than a month of Premium at $79.99.

So the phone is where the volume goes. The render is where the scripts nobody will film go.

That split is the whole plan. It does not depend on anybody's opinion about how convincing renders look.

It also explains the shape of the spend. Nobody needs twenty announcements a month. Everybody needs twenty attempts.

One account covers both. That is why this comparison does not end with you buying a second tool.

When to send the script the other way

Nobody in the building will speak on camera or off it, and there is no budget for a creator.

The script is a repeating announcement where identical delivery across months is the point.

You are testing six openings and want the delivery held constant, so the words are the only variable.

The person who would film is unavailable and the ad has to ship this week.

In each case, render it at 350 credits a minute. Batch the render exactly as you would a filmed take.

The plain facts on both routes. A filmed source take caps at three minutes. Rendered audio caps at five. Output is 9:16.

There is no timeline, by design. Every lever is a decision about the ad.

Filming has one cost the price does not show. Somebody has to be asked, scheduled and willing again next month.

Teams that film sustainably batch four scripts into one three minute session. That habit is worth more than any setting in the product.

Whichever route the script takes, the file that comes back is finished. That is the part you would otherwise be doing in a second tool.

Questions people ask

I am bad on camera. Does that change the maths?
Less than you expect, because the take is not the ad. Three minutes of imperfect talking gets cut to forty seconds and you delete the lines you dislike. A bad take and a good one cost the same 100 credits.
What equipment do I need?
A phone, a windowsill and somewhere to prop it at eye height. Stand side on to daylight, two paces off the wall. That is the whole setup.
Can I film and still keep my face out of the ad?
Yes. Point the camera at your hands and the product while you talk. The batch reads the transcript from your speech. The ad works identically and no face appears.
Who should buy the rendered route instead?
Teams whose ads are recurring announcements, and teams where nobody will speak at all. Use both in the same account rather than choosing once.

Three minutes, one phone, a 100 credit batch and 20 credits a minute to export, and a vertical ad comes back with captions and b-roll already on it.

3 videos free, no card3 finished videos free in your first 7 days, no card. They carry a Cutroom mark; Lite at $19.99/month removes it