Feature
Voice-over for avatar video: three routes, one of them free
The audio decides how a render looks. Three routes feed it. Generate a read at 70 credits a minute, record your own for nothing, or hire somebody. Here is what each route does to the sync, what each costs against the render it feeds, and how to pick one in under a minute.

- Per minute of generated voice-over
- 70 CR
- Recording your own, in a soft room
- 0 CR
- Audio window the render accepts
- 3s to 5min
Route one: generate it, 70 credits a minute
A generated read is recorded in no room at all. No reverb, no fridge hum, no bitrate history.
That makes it the most reliable input for lip sync. The waveform the model reads is clean.
Forty seconds is about 47 credits. That is small next to the 233 the render costs at that length.
It is also the only route that needs nobody's time. Write, generate, render, finished, without leaving the desk.
The cost is delivery. A generated read is even. Evenness is what makes a longer piece feel synthetic.
Under forty seconds that is barely noticeable. Over ninety it is the main thing a viewer hears.
So it is the right pick for announcements, offers, spec walkthroughs and anything where the information is the value.
For those scripts an even read is not a weakness. It is the delivery the sentence wanted.
Voice-over cost for a forty second ad
The voice is the cheap part. What it costs you is delivery, not credits.
- Record your own0 CR
- Generated readAbout 47 CR
- The render it feedsAbout 233 CR
Route two: record your own, free, and mind the room
Your own voice costs nothing and carries variation no generated read has.
It also carries whatever room you recorded it in.
Reverberant spaces are the enemy. A kitchen, a bathroom, a bare office.
The tail of each word smears into the next and the mouth never fully closes.
Record somewhere soft. Carpet, curtains, a wardrobe of clothes behind you.
Give it speech alone. No music, no ambience, no ducking applied before upload. The model cannot separate a bed from a voice.
Export at a high bitrate. Consonants are the highest frequency information in the file and the first casualty of a low one.
Aim near 150 words a minute. Never above 180, where mouth shapes start to overlap.
Done that way, your own voice on your own face is the strongest version of this format.
- Soft room, speech only, high bitrate, near 150 words a minute
- Half a second of silence at the top and tail
- Listen to it alone before rendering. Unclear to the ear means unclear on the mouth.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Route three: hire a voice, and know what it buys
A hired voice gives you performance. Emphasis, pace changes, a pause where a thought lands.
That is worth paying for on a script where one specific line has to hit.
It is worth nothing on an announcement. An even read does the same job there for 70 credits a minute.
The practical catch is turnaround. A render takes minutes. A booking takes days.
Ask for a dry read with no processing, no music and no room reverb. Studio polish added at their end can hurt the sync.
Ask for the raw file at a high bitrate rather than a mastered mp3.
For most teams this route makes sense on one hero ad a quarter and on nothing else.
The other two routes cover the rest of the calendar without anybody being booked.
Picking a route in four questions
Most scripts answer this in under a minute.
Does a line need emphasis?
Yes: your voice or a hired one
Is there a soft room?
No: generate it
Is it under 40 seconds?
Yes: generated is fine
Is it testimony?
Yes: film it instead
What every route shares: the audio decides the mouth
Nothing about the portrait tells the model what shape the mouth should make. The waveform does all of it.
That is why a clean generated read often syncs better than a real voice recorded badly.
It also means every fix here is free before it is paid. Listen alone, check the pace, check the room.
The audio window is three seconds to five minutes. Below three it is refused. Above five it is refused.
At 350 credits per rendered minute, the range most people use is twenty to forty-five seconds.
Test a twenty second version first. About 117 credits buys the answer to whether the voice and the face agree.
Run that test even when you are confident. Voice and portrait pairings fail in ways neither input predicts alone.
A deep voice on a young face lands wrong. So does a fast read on a still expression. Each looks fine on its own.
Keep the voice-over. It is an asset, not an input.
Save the audio file after the render, because you will want it again.
The same read can front a second portrait, or sit under product footage with no presenter at all.
Neither of those costs another 70 credits a minute.
In the editor you highlight a phrase and a clip lands over exactly those words. The same voice can carry three different ads.
That is how teams stop paying twice for the same forty seconds of speech.
Where a voice cannot help is testimony and sensory claims. A hired read over a rendered face is still a face that felt nothing.
Film those. Three minutes on a phone batches for 100 credits. That brings the voice and the face together.
The plain facts. Output is 9:16. There is no timeline, by design.
And the permission rule stands whoever is speaking. The portrait is yours or belongs to somebody who gave explicit written permission.
One account holds the voice, the render and the cut. That is why the same read keeps earning after the first ad.
Questions people ask
- Does a generated voice sync better than my own?
- Usually, unless you have a soft room. A generated read has no reverb, no ambience and no compression history, which is exactly what the sync is reading.
- Can I use a voice from a video I already made?
- If you can extract clean speech with no music or ambience under it, yes. If a bed is mixed in, the model gets two signals and the sync drifts.
- How much does the voice change the finished cost?
- Little. At forty seconds a generated read is about 47 credits against roughly 233 for the render and 100 for the batch. It is the cheapest decision with the most effect on believability.
- Who should not be choosing a voice-over route?
- Anybody whose script is testimony or demonstration. No voice rescues those from a rendered face. Film a three minute take and batch it for 100 credits instead.
Clean speech, a soft room or a generated read, near 150 words a minute. Then one batch turns the render into an ad.