Skip to content

Feature

How AI avatars work: the photo is the least important input

An avatar render has three inputs. People rank them in the wrong order. The audio drives the mouth. The script decides whether anybody believes it. The portrait mostly rules options out. Here is the pipeline priced step by step, what each input controls, and the step after the render that turns a take into an ad.

A man with glasses focused on editing a photo on his desktop computer in an office.
Photo by Malte Luk on Pexels
Still image the render is built from
1
Audio window the render accepts
3s to 5min
Per finished minute of render
350 CR

The pipeline in five steps, with a price on each

You write a script. That step is free and it decides more of the outcome than everything after it combined.

The script becomes audio, either recorded by you or generated here at 70 credits a minute.

The audio and a portrait go to the render. Back comes a 9:16 MP4 of that face speaking. It bills at 350 credits per finished minute.

That MP4 is a source take. Batching it into a real ad is 100 credits. It returns a transcript-driven cut with captions and b-roll.

Export is 20 credits per output minute. Forty seconds finished lands near 393 credits from blank page to file.

Nothing in that chain is mysterious. Every step has a number on it.

Knowing the numbers stops the expensive mistake. That mistake is spending 525 credits animating ninety seconds nobody will watch.

The first three videos are free, over 7 days, with no card. Walk the whole chain once before paying for any of it.

What happens between your script and the MP4

Five steps. The free one decides the most.

  1. Script

    Free. Read it out loud.

  2. Audio

    Recorded, or 70 CR a minute

  3. Render

    350 CR a minute, lip-synced

  4. Batch

    100 CR, cut and captions

  5. Export

    20 CR per output minute

The audio is what the mouth is reading

Lip movement is derived from the waveform of the speech you supplied.

Nothing about the portrait tells the model what shape the mouth should make.

That is why a reverberant recording produces mushy mouth shapes. The tail of each word smears into the next. The model follows it faithfully.

Music or ambience mixed under the voice before upload makes it worse. The model cannot tell which part of the sound is speech.

A low bitrate export removes consonants. Consonants are the highest frequency information in the file and the sharpest cues the mouth has.

A phone recorded in a room full of soft furnishings beats a good microphone in a bare office. Every time.

A generated read sidesteps all of it for 70 credits a minute. It was recorded in no room at all.

Listen to the audio alone before you spend a credit. If it sounds unclear to your ear, it will look unclear on the mouth.

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

The portrait constrains the render rather than improving it

The model works out how a whole face moves from one frozen instant. Ambiguity in the still becomes artefact in the video.

A hard shadow across one cheek gets baked in and then has to move. That is why it reads as slightly wrong.

Portrait-mode blur smears once the head starts moving. The blurred region was never a real depth map.

A wide smile is the single most common cause of a strange result. The video has to travel away from that expression.

A three-quarter angle photographs better and renders worse. Less of the far side of the face is visible to work from.

None of these mean a better photo produces a better video. They mean a worse photo produces a worse one.

That is a different relationship. It is why ten minutes by a window settles the portrait for good.

Shoot five. Render a test from the best two. Stop optimising the picture after that.

  • Front on, eyes to the lens, shoulders in frame
  • Even soft light, plain wall behind, nothing crossing the face
  • Mouth closed or barely parted, expression neutral
  • PNG, JPEG or WebP under 15 MB, unprocessed and unfiltered

How much each input decides

Most buyers spend the afternoon on the third one. The order is worth reversing.

  • The scriptMost of it
  • The audio recordingThe next slice
  • The portraitRules options out

The render produces a take, and the take is the wrong shape for a feed

What comes back is one continuous shot of a head and shoulders. No cutaway, no second angle, no movement through a room.

That format is the most skippable thing in a vertical feed, whether the face is generated or filmed.

The batch step fixes it. Most pipelines stop one step earlier and hand you the file.

Here you direct the cut on the transcript rather than on a timeline.

Highlight a phrase and a clip lands over exactly those words. Delete a line and the cut rebuilds around it.

Change the pace and the whole piece re-cuts. Chill covers up to 40 percent of the runtime with footage, normal 45, fast 52.

Captions come from six style packs, and each clip enters with a cut, a whip, a punch, a glitch or a sweep.

Music sits under it for 50 credits a track.

Skipping this step is the most common reason somebody tries the category once and decides it does not work.

What the pipeline does not do, and where that leaves you

It does not produce hands. No stage of the render puts a limb or a product in the frame.

It does not produce spontaneity. The half laugh, the restarted sentence and the thinking pause are not in its vocabulary.

It does not make a claim personal. A face saying it changed my mornings is performing a feeling it never had.

It does not exceed five minutes of audio. At 350 credits a minute nobody sensible goes near that ceiling.

And it cannot be pointed at somebody who did not agree. The portrait is yours or belongs to a person who gave explicit written permission.

The plain facts alongside those. Output is 9:16. There is no timeline, by design.

What that leaves is a clean division of labour. Render the sentences that carry information. Film the sentences that carry a person.

Both arrive in the same editor. Both leave as the same finished vertical file.

A render pipeline on its own has never given anyone that second half.

Questions people ask

Does the model animate the whole photo or only the mouth?
The mouth and the immediate face region move, with small head motion around it. That is why rigid edges near the jaw, like a high collar or a hat brim, read as detached.
Why does my render look worse than the demo reels?
Almost always the audio or the script. Demo reels use studio-clean speech and sentences written to be spoken. A kitchen recording of writing meant for the page renders as poorly as it reads aloud.
Can I improve a render after it comes back?
Not the render itself. You improve the ad around it. Batch it for 100 credits, cover phrases with footage, add captions and cut the length down. That is where most of the quality was always going to come from.
Who does not need this pipeline at all?
Anybody with a willing person and a phone whose scripts are all testimony. Film those takes and batch them for 100 credits in the same editor.

Script, audio, portrait, then the batch that turns the take into an ad. That fourth step is the one most pipelines leave out.

3 videos free, no card3 finished videos free in your first 7 days, no card. They carry a Cutroom mark; Lite at $19.99/month removes it