Feature

One photo, one audio track, and a presenter renders

Two files decide everything here. One front-facing photograph and one audio track return a lip-synced presenter take at 320 credits a rendered minute. That take then runs through the editor exactly like filmed footage and exports as a finished 9:16 ad. Here is what makes a photo work, what the audio decides, and what a campaign costs. By the end you will know which scripts to render and which to film.

A hand holding a smartphone showing an incoming call notification with a photo of a man on the screen.
Photo by cottonbro studio on Pexels
Photo required
1
Credits per avatar minute
320
Credits per voice-over minute
70

What you hand it, what it returns, and what the minute costs

The inputs are two. One still photograph of a person, front-facing, with nothing crossing the mouth or jaw. One audio track of the read.

The output is a video take of that person speaking your words. It enters the editor exactly like a filmed one. Transcribed, cut, covered with footage on the words that need showing, captioned, headlined, exported 9:16.

A rendered minute is 320 credits. That makes it the most expensive thing in the product by a wide margin, because every frame is generated rather than found.

A forty second read is about 213 credits of render. The 100 credit batch and the export at 20 credits a minute sit on top of it.

Basic at $39.99 for 2,500 credits a month is a little under eight minutes of presenter footage. Premium at $79.99 for 5,000 is about fifteen. A $15 top-up adds 1,000 credits, roughly three more minutes.

Credits per minute, by what the machine makes

Generating frames costs sixteen times what exporting them does. Write the avatar read short.

  • Export a finished minute20 CR
  • Generated voice-over, a minute70 CR
  • One batch, whole take100 CR
  • Avatar render, a minute320 CR

It animates a photograph in time with a waveform, and that predicts everything

The model does one narrow task. It drives a mouth, a jaw and some head motion from an audio waveform. The face it works from is a single still frame.

Hold that sentence and you can predict the result before spending a credit. Anything near the mouth tends to look right, because that is the whole job.

Anything far from the mouth is where the illusion goes thin. Hands, shoulders, the background, the sense that a body is holding itself up. None of it was ever being modelled.

Which is why the practical advice is short reads and frequent cuts to footage. The illusion is strongest in bursts and thins over long holds.

What the model is handed, and what it drives

Nothing outside the mouth and jaw is being modelled, which is exactly where viewers look next.

  1. One still photo

    Front-facing, mouth clear

  2. One audio track

    Yours, or 70 CR a minute

  3. Mouth, jaw, head

    Driven from the waveform

  4. A video take

    320 CR per rendered minute

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

The photo sets the ceiling, the audio decides whether you reach it

Front-facing beats three-quarter. A neutral or slightly open mouth beats a wide grin, because a fixed smile has to be broken apart to speak. Even, soft light beats hard shadow.

Nothing may cross the mouth or jaw. No hand on the chin, no microphone, no heavy shadow at the lip line.

Resolution matters less than people expect. Framing matters far more. A clean head-and-shoulders shot at moderate resolution beats a high-resolution full-body photo where the face is a hundred pixels wide.

The audio carries more of the result than the image does. Clear speech, even pace, no traffic behind it, no music baked in. Muddy audio gives the mouth nothing crisp to follow and the sync goes soft.

Record the read yourself where you can. It usually sounds better and it costs nothing extra.

Test the photo before you write the script. One thirty second render is 160 credits, and it tells you more than any checklist can.

The presenter carries the words. Footage carries what it cannot hold

The render gives you speech and modest head movement. That is the performance in full, and for an information-led script it is enough.

There are no hands. The presenter cannot pick anything up, point at anything, or demonstrate a step.

So you cover those lines with real footage of the product, placed on the exact words that name it. Highlight the phrase and the shot lands there. It does the job a gesture would have done, and does it more clearly.

There is no acting either. No glance to camera at the punchline, no shift in posture when the tone changes. Write for the voice and let the pictures do the pointing.

One photo means one look for the length of the take. A two minute render holds the same expression at the end as at the start, which is why short reads and frequent cuts win.

And the face has to be one you may use. Your own, or somebody who agreed in writing that their likeness can appear in your paid advertising.

Where a render beats a booking, and where a phone still wins

The render earns its price on three things. Nobody has to be available. Ad twenty is framed exactly like ad one. Twenty scripts can be rendered inside a week.

That is a strong case. A spokesperson effect that used to need twenty bookings now needs one photograph and one audio file.

A filmed take is cheaper. One batch is 100 credits against 320 a minute of render. It can hold the product, and audiences believe it more.

So the split is clean. Render the scripts nobody will read on camera. Film the ones where a specific human has to be believed. Both routes use the same editor.

A medical result or a financial outcome belongs in a filmed take. Some viewers clock a render on a long hold, and the claim inherits the doubt.

What is left for the render is large. The ad carried by what is said. A founder who will not be on camera. A presenter you want in twenty ads over six months without booking anybody twice.

A generated presenter against filming it yourself

Two rows go to the camera. Render the five it wins, and film the two it does not.

Filming a real takeCutroom
Holding or pointing at the productHands are in frame, doing the sellingNo hands, no props, no gestures
Being believed on a trust-led claimA real person carries the claimSome viewers clock it on longer holds
Ad twenty framed exactly like ad oneLight and wardrobe drift over six monthsSame photo, same framing, every render
Works when nobody will go on cameraSomebody has to do the readA photo and an audio track is enough
Twenty scripts turned around in a weekTwenty set-ups and twenty diaries320 credits per rendered minute
Changing the offer the afternoon it changesBook, light and film it againNew audio, same photo, re-render
Captions and footage on the same fileA second tool after the shootThe render goes straight into the editor

Questions people ask

How many photos do I need?
One. A front-facing head-and-shoulders shot in even light, with nothing crossing the mouth or jaw, beats a whole folder of badly framed images.
Can I use my own voice?
Yes, and it usually sounds better. Record the read and supply it as the audio. If you would rather the machine spoke it, generated voice-over costs 70 credits per minute.
What does a minute of avatar cost?
320 credits. Basic at $39.99 for 2,500 credits a month is a little under eight minutes, Premium at $79.99 for 5,000 is about fifteen, and a $15 top-up adds 1,000 credits.
Can I use a photo of a public figure or a stock model?
No. Use yourself, a colleague who has agreed in writing, or a licensed likeness where the licence explicitly covers synthetic video. Public figures are never available.
Is it right for every script?
It is built for scripts where the information is the value. If the ad has to show the product being handled, film that part on a phone and run it through the same editor for 100 credits, then place it under the presenter's words.

One photo, one audio track, and the same presenter in ad twenty. Then footage on the words a face cannot hold, in the same tool.

Start with one take300 free credits · no card · cancel anytime