Guide
An avatar video is fronted by a generated presenter. Nobody stood in front of a camera
An avatar video is a video fronted by a generated presenter instead of a filmed person. The presenter is built from a photograph and a voice track, with lip movement synced to the words. Nobody stood in front of a camera, and everything else about the video is normal. This page covers how one is made, where it convinces and where it does not.
- Photograph the presenter is built from
- 1
- Credits per rendered minute in Cutroom
- 350
- Lights, retakes or filming days involved
- 0
A photograph and a voice track go in, a presenter comes out
The inputs are a still image of a face and an audio track of the script. The output is that face delivering the script, with the mouth, head and expression moving through the read.
Modern systems render the performance as one continuous generation rather than looping a short clip. That is what stopped avatar video looking like a puppet with three mouth shapes.
The voice track can be a recording of a real person or a generated voice reading the script. Either way the audio comes first, and the face is animated to match it.
How an avatar video is assembled
The audio leads. The face is animated to the voice, never the other way around.
One photograph
A clear, front-facing face
A voice track
Recorded or generated
The render
Continuous, lip-synced
A normal edit
Captions, b-roll, music
Where it works, and where a viewer catches it
Informational delivery works. A presenter explaining a feature, walking through steps or reading an announcement reads as a presenter, and the format holds.
First-person testimony does not. A generated face claiming this product changed my life asks the viewer to believe a feeling the face never had, and viewers are unusually good at catching that.
The failure is social, not technical. More resolution does not fix it, because the problem is the claim, not the pixels. Write avatars informational scripts and film humans for testimony.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
What it replaces, and what it costs
The economics replace a filming day, not an editor. No lighting, no wardrobe, no seventh take because a phone buzzed, and a script revision is a re-render rather than a re-shoot.
In Cutroom an avatar render is priced by output length at 350 credits per rendered minute, and the result drops into the same pipeline as filmed footage.
The honest comparison is against your own face on camera, which is free to film and better at being you. Avatars earn their cost when filming is the bottleneck, not when it is merely unglamorous.
Questions people ask
- Is an avatar video the same as a deepfake?
- The underlying technology is related, the use is not. An avatar video uses a face with the owner's consent to deliver a script openly. A deepfake puts words or acts on someone without consent. Consent and disclosure are the line.
- Should I disclose that the presenter is generated?
- Yes. Several platforms already require labelling AI-generated realistic people, and the rules are tightening rather than loosening. Disclosure also protects the trust the video is trying to build.
- Can the avatar be my own face?
- Yes, from a photograph of you, and that is the strongest use. Your face, your voice track, and none of your filming afternoons.
An avatar removes the camera from the process, not the judgement. Give it the informational scripts, keep the testimony human, and label it plainly.