Feature
One photo, one audio track, and a presenter renders
One front-facing photo and one audio track return a lip-synced presenter take at 350 credits a rendered minute. That take runs through the editor like filmed footage and exports as a finished 9:16 ad. Here is what makes a photo work, what the audio decides, and what a campaign costs.

- Photo required
- 1
- Credits per avatar minute
- 350
- Credits per voice-over minute
- 70
What you hand it, what it returns, and what the minute costs
The inputs are two. One still photograph of a person, front-facing, with nothing crossing the mouth or jaw. One audio track of the read.
The output is a video take of that person speaking your words. It enters the editor exactly like a filmed one. Transcribed, cut, covered with footage on the words that need showing, captioned, headlined, exported 9:16.
A rendered minute is 350 credits. That makes it the most expensive thing in the product by a wide margin, because every frame is generated rather than found.
A forty second read is about 233 credits of render. The 100 credit batch and the export at 20 credits a minute sit on top of it.
Basic at $39.99 for 2,500 credits a month is a little over seven minutes of presenter footage. Premium at $79.99 for 5,000 is about fourteen. A $39.99 top-up adds 2,500 credits, roughly seven more minutes.
Credits per minute, by what the machine makes
Generating frames costs sixteen times what exporting them does. Write the avatar read short.
- Export a finished minute20 CR
- Generated voice-over, a minute70 CR
- One batch, whole take100 CR
- Avatar render, a minute350 CR
It animates a photograph in time with a waveform, and that predicts everything
The model does one narrow task. It drives a mouth, a jaw and some head motion from an audio waveform. The face it works from is a single still frame.
Hold that sentence and you can predict the result before spending a credit. Anything near the mouth tends to look right, because that is the whole job.
Anything far from the mouth is where the illusion goes thin. Hands, shoulders, the background, the sense that a body is holding itself up. None of it was ever being modelled.
Which is why the practical advice is short reads and frequent cuts to footage. The illusion is strongest in bursts and thins over long holds.
What the model is handed, and what it drives
Nothing outside the mouth and jaw is being modelled, which is exactly where viewers look next.
One still photo
Front-facing, mouth clear
One audio track
Yours, or 70 CR a minute
Mouth, jaw, head
Driven from the waveform
A video take
350 CR per rendered minute
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
The photo sets the ceiling, the audio decides whether you reach it
Front-facing beats three-quarter. A neutral or slightly open mouth beats a wide grin, because a fixed smile has to be broken apart to speak. Even, soft light beats hard shadow.
Nothing may cross the mouth or jaw. No hand on the chin, no microphone, no heavy shadow at the lip line.
Resolution matters less than people expect. Framing matters far more. A clean head-and-shoulders shot at moderate resolution beats a high-resolution full-body photo where the face is a hundred pixels wide.
The audio carries more of the result than the image does. Clear speech, even pace, no traffic behind it, no music baked in. Muddy audio gives the mouth nothing crisp to follow and the sync goes soft.
Record the read yourself where you can. It usually sounds better and it costs nothing extra.
Test the photo before you write the script. One thirty second render is 175 credits, and it tells you more than any checklist can.
The presenter carries the words. Footage carries what it cannot hold
The render gives you speech and modest head movement. That is the performance in full, and for an information-led script it is enough.
There are no hands. The presenter cannot pick anything up, point at anything, or demonstrate a step.
So you cover those lines with real footage of the product, placed on the exact words that name it. Highlight the phrase and the shot lands there. It does the job a gesture would have done, and does it more clearly.
There is no acting either. No glance to camera at the punchline, no shift in posture when the tone changes. Write for the voice and let the pictures do the pointing.
One photo means one look for the length of the take. A two minute render holds the same expression at the end as at the start, which is why short reads and frequent cuts win.
And the face has to be one you may use. Your own, or somebody who agreed in writing that their likeness can appear in your paid advertising.
Where a render beats a booking, and where a phone still wins
The render earns its price on three things. Nobody has to be available. Ad twenty is framed exactly like ad one. Twenty scripts can be rendered inside a week.
That is a strong case. A spokesperson effect that used to need twenty bookings now needs one photograph and one audio file.
A filmed take is cheaper. One batch is 100 credits against 350 a minute of render. It can hold the product, and audiences believe it more.
So the split is clean. Render the scripts nobody will read on camera. Film the ones where a specific human has to be believed. Both routes use the same editor.
A medical result or a financial outcome belongs in a filmed take. Some viewers clock a render on a long hold, and the claim inherits the doubt.
What is left for the render is large. The ad carried by what is said. A founder who will not be on camera. A presenter you want in twenty ads over six months without booking anybody twice.
A generated presenter against filming it yourself
Two rows go to the camera. Render the five it wins, and film the two it does not.
| Filming a real take | Cutroom | |
|---|---|---|
| Holding or pointing at the product | Hands are in frame, doing the selling | No hands, no props, no gestures |
| Being believed on a trust-led claim | A real person carries the claim | Some viewers clock it on longer holds |
| Ad twenty framed exactly like ad one | Light and wardrobe drift over six months | Same photo, same framing, every render |
| Works when nobody will go on camera | Somebody has to do the read | A photo and an audio track is enough |
| Twenty scripts turned around in a week | Twenty set-ups and twenty diaries | 350 credits per rendered minute |
| Changing the offer the afternoon it changes | Book, light and film it again | New audio, same photo, re-render |
| Captions and footage on the same file | A second tool after the shoot | The render goes straight into the editor |
Questions people ask
- How many photos do I need?
- One. A front-facing head-and-shoulders shot in even light, with nothing crossing the mouth or jaw, beats a whole folder of badly framed images.
- Can I use my own voice?
- Yes, and it usually sounds better. Record the read and supply it as the audio. If you would rather the machine spoke it, generated voice-over costs 70 credits per minute.
- What does a minute of avatar cost?
- 350 credits. Basic at $39.99 for 2,500 credits a month is a little under eight minutes, Premium at $79.99 for 5,000 is about fifteen, and a $39.99 top-up adds 2,500 credits.
- Can I use a photo of a public figure or a stock model?
- No. Use yourself, a colleague who has agreed in writing, or a licensed likeness where the licence explicitly covers synthetic video. Public figures are never available.
- Is it right for every script?
- It is built for scripts where the information is the value. If the ad has to show the product being handled, film that part on a phone and run it through the same editor for 100 credits, then place it under the presenter's words.
One photo, one audio track, and the same presenter in ad twenty. Then footage on the words a face cannot hold, in the same tool.