Feature

Talking head generator: what goes in, what comes back, what it costs

Talking head generators all do the same first thing well. A portrait and an audio file return a video of that person speaking. The mouth is in time with the sound. Almost none of them do the second thing. One unbroken shot of a face is not an ad, and that part lands on you. Here is what the render needs, what a minute costs, and how Cutroom finishes the file.

Asian woman smiling while talking on a phone in a modern office setting.
Photo by Mikhail Nilov on Pexels
One batch: a take becomes a finished ad
100 CR
Per minute of generated presenter
320 CR
Longest filmed take the uploader accepts
3 min

Two files in, a lip-synced vertical MP4 out

Give it one portrait and one audio track. That is the whole input.

The portrait is a PNG, JPEG or WebP under 15 MB. Front on, evenly lit, mouth closed.

The audio runs from three seconds to five minutes. Record it yourself, or generate a read here for 70 credits a minute.

Back comes a vertical MP4 of that face speaking those words. Head and shoulders, one shot. The camera does not move.

There are no hands, no set change and no second angle. That is the shape of the format, not a setting you missed.

The permission rule comes before anything else. The face is yours, or belongs to somebody who gave explicit written permission.

The 300 trial credits buy about a minute of render. No card. The whole route is testable before you pay.

A forty second clip with a generated read lands near 260 credits all in.

  • Photo: PNG, JPEG or WebP under 15 MB, face to the lens
  • Audio: three seconds to five minutes, speech alone
  • Out: a lip-synced 9:16 MP4, chest up, one continuous shot

What a minute costs, next to the filmed route in the same account

Generated presenter is 320 credits per finished minute. Forty seconds is about 213. Add 47 if the read is generated.

Cutroom also takes a filmed take. Upload one talking-head recording of up to three minutes. One batch of 100 credits turns it into a finished ad.

That batch covers the transcription, the director pass and the b-roll search. Export is 20 credits per output minute on top.

So a three minute source take cut to forty seconds is 100 plus about 14 credits.

Both files land in the same editor. Both come out as the same 9:16 MP4.

Send the render the scripts where the information is the value. Announcements, offers, spec walkthroughs, anything nobody wants to say five times.

Send the phone the claims that only mean something because a person made them. I have used this since March is one of those.

You never have to choose once. Both routes sit in one account, and a single ad can use both.

Credits for the same forty second ad

Both routes end in a 9:16 MP4 from the same editor. Pick by what the script needs.

  • Filmed take, batched and exported114 CR
  • Generated presenter with generated read260 CR

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

One batch turns either take into a finished ad

Both routes hand you one continuous shot of a person speaking. That file underperforms in a feed either way.

The editor works on the transcript. You read the words and make decisions about the ad instead of about software.

Highlight a phrase and a clip lands over exactly those words. Delete a line you do not want and the cut rebuilds around it.

Change the pace and the machine re-cuts the whole piece. Chill covers up to 40 percent of the runtime with footage, normal 45, fast 52.

Emphasise a word so it pops in the captions. Pick from six caption packs. Set the colour, weight, size and position.

Choose how each clip enters with a cut, a whip, a punch, a glitch or a sweep.

Swap the clip in any spot, or search millions of free ones. Turn music on for 20 credits and set how loud it sits.

There are no keyframes and no layers. That is the point. Every lever is a decision about the ad.

This is the step most generators leave to you. It is why the file you download here is uploadable.

What one batch does to a take

The same five steps run whether the take was filmed or rendered.

  1. Take in

    Filmed, up to 3 minutes, or rendered

  2. Transcript

    Every word, editable

  3. Director pass

    Cut, hook, caption plan

  4. Your edits

    Highlight, delete, re-pace

  5. 9:16 MP4

    20 CR per output minute

The four problems you will hit in the first week

Reverberant audio. Lip movement follows the waveform. A kitchen recording gives mushy mouth shapes that read as the software failing.

The fix is a soft room. Carpet, curtains, a wardrobe of clothes behind you. Or generate the read. It was recorded in no room at all.

Scripts written for the eye. A semicolon and a subordinate clause sound wrong in any mouth. Say it out loud before you upload it.

The wrong portrait. A wide smile, a hard side light or a portrait-mode blur will each quietly cost you a render.

Length. At 320 credits a minute, a ninety second script is 480 credits of render. It holds nobody anyway.

Forty seconds is the working ceiling. One point, delivered, then out.

A twenty second test render catches every one of them for about 107 credits.

Run that test before you write the rest of the script.

Where the render earns its 320 credits a minute

Announcements. Pricing changes, shipping cutoffs, restock notices, policy updates, spec comparisons.

Nobody volunteers to read those on camera. Nobody watching cares who reads them. The information is the value.

Series consistency. The same presenter fronts week one and week twenty. Nobody gets rebooked or restyled.

Identical delivery across a test. When you are trying six openings, the delivery should be the one thing holding still.

A person films six takes and gives you six energies. Six lightings too. The winner you find may be a mood rather than a message.

The plain facts to plan around. There are no hands in a rendered frame. Anything held up to the lens gets filmed and placed under the words.

Output is 9:16. Rendered audio caps at five minutes. A filmed source take caps at three.

There is no timeline, by design. Every lever is a decision about the ad.

None of that touches what the format is good at. None of it sends you to a second tool either.

Questions people ask

Which route should I start with?
Whichever the script demands. Announcements and explainers go to the render. A claim that only means something because a person made it goes to the phone. Both end in the same editor.
Can I mix them?
Yes, and it is often the best answer. Generate the presenter for the explanatory middle. Film the opening line on a phone. The editor puts b-roll over the phrases that need showing.
How long should the finished ad be?
Forty seconds or under for a generated presenter, because an even delivery has no small human variation to earn patience with. A filmed take can hold longer.
Who is this the wrong purchase for?
Anybody who needs a horizontal 16:9 file or frame-level control. There is no timeline here by design. A conventional editor is the right buy for that work.

Two ways in, one 9:16 MP4 out, captions and b-roll already on it. Most generators hand you the clip and wish you luck.

Start with one take300 free credits · no card · cancel anytime