Comparison

HeyGen delivers a presenter. Cutroom delivers the advert

HeyGen renders a consistent presenter reading your script. Many languages, any length, no reshoot when a price moves. What it hands back is one unbroken shot of a face. That is a clip, not an advert. This page covers what an advert needs on top, what one costs in credits, and what a first week of recorded takes looks like. By the end you will know which of the two jobs your week is actually short of.

Professional presentation setting with a screen and laptop, captured indoors.
Photo by Matheus Bertelli on Pexels
The only shape Cutroom exports
9:16
Hard ceiling on a source take
3 min
Caption packs shipped, no pack builder
6

HeyGen is right about the presenter. The advert is the next job.

Three properties a transcript-first ad tool does not attempt.

Consistency across a large library. One presenter, one voice, two hundred videos recorded over a year, with no drift in haircut, lighting or mood.

Language reach. The same script delivered in many languages is a hard problem. It is a distribution multiplier if you sell across borders.

Correction without reshooting. A price changes. You change the line. You re-render. Anyone who has re-filmed forty seconds because a number moved knows what that is worth.

None of that is gloss. Their tiers are published on their own site and they move, so read them there on the day.

Now the part that stays on your evening. Forty seconds of a face talking, no cuts, no cutaways, no captions, is not an advert. That is true of filmed footage too. Better rendering does not fix it. Cutroom does that step in one more move.

HeyGen and Cutroom, row by row

Two rows go to HeyGen outright, and they are the two worth knowing. The other five are what a feed asks for after the presenter is rendered.

HeyGenCutroom
The same script in several languagesA distribution multiplierA take is in one language
Video longer than three minutesLessons, onboarding, docsThree minute source ceiling
The cut rather than the deliveryHands you a person talkingTrim, cutaways, captions, pace
Cutaways on the exact phraseOne unbroken shotHighlight, and it lands there
Coverage held under a ceilingNot a constraint it enforces40, 45 or 52 percent by pace
A 9:16 file you can upload todayA clip, then your own editCut, captioned, exported
A viewer who cannot tell it was producedClean reads as an advertA real room, a real pause

A presenter video is information delivered. A feed ad is a fight for second three.

Those are different formats with different requirements. The difference is invisible from inside the tool that made the video.

A vertical ad asks for four decisions a presenter render never makes. A hook that earns the third second. Cutaways that reset attention before it drifts. Captions carrying the message, because a large share of the audience has the sound off. A pace that changes when the argument changes.

Every one of those is an editing decision. A presenter generator does not make them. Making them was never its job.

So the usual patch is exporting the file into a separate editor. That works. It is also how a fifteen minute task becomes ninety, twelve times a month, forever.

Cutroom collapses the patch. One take in. One batch at 100 credits. The cut comes back made.

Then you direct it on the transcript. Highlight a phrase and a clip lands over exactly those words. Delete a line and the cut rebuilds. Change the pace and the whole read re-cuts.

How much of the runtime a cutaway may cover

Enforced by the compiler, not suggested in a help article. The cap is why a recorded take still reads as a person talking rather than a stock montage with narration.

  • Chill pace40%
  • Normal pace45%
  • Fast pace52%

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

Writing for the ear moves with you, and it is the expensive part

Generated-presenter work builds one skill fast. Anything clumsy is immediately audible when a machine reads it back flat. That skill moves with you and it is worth more than people credit.

So does your hook library, your testing cadence and your sense of what a delivered line has to sound like.

Two facts are worth stating flatly before you plan a month. A take is in one language. Three minutes is the ceiling on a source take, so a twelve minute onboarding module belongs elsewhere.

Neither is a plan tier you can buy past. Both are the shape of a short-form ad tool.

  • Moves with you: scripts, your hook library, the habit of writing for spoken delivery, your testing cadence.
  • Moves in a different shape: the presenter idea. A photo plus a voice-over becomes a lip-synced presenter at 320 credits per rendered minute.
  • Left behind: multi-language output, video over three minutes, and presenters you did not supply.
  • Left behind: horizontal and square exports. Everything here is a 9:16 MP4.
  • Gained: the cut. Footage on the words you highlight, captions rendered in, pace as one lever, a finished file rather than a clip to assemble.

Record three scripts near a window, then ask which three seconds you would change

Pick the three scripts that already perform best. Record each as one continuous take on a phone. Stand near a window. No lighting kit. Under three minutes.

This step feels like a downgrade. The feeling is the point. The small imperfections you are worried about are what makes a file read as native in a feed rather than as an advertisement.

Run one batch at 100 credits. Look at the draft with a specific question. Not is it good. Which three seconds would I change.

The draft places footage under the phrases where you named something concrete. Those are the words with a picture waiting behind them.

Fix the misses by highlighting different words. Set the pace and watch the coverage cap hold at 40 percent for chill, 45 for normal, 52 for fast.

Export at 20 credits per output minute. Run it against your best current creative in the same ad set for a week. The whole test is about 110 credits and half an hour, inside a trial of 300 credits over 7 days with no card.

Run both, and let each one do the job it was built for

The clean split is by format rather than by loyalty. A lesson is information for somebody who already chose to watch. An ad is a fight for the next two seconds with somebody who did not.

Keep a presenter platform for the library. Localised modules, onboarding, documentation, anything over three minutes, one face across two hundred assets. Those are its jobs and it does them well.

Put the paid ads here. Three facts to plan around. Three minutes of source per take. 9:16 MP4 out. Six caption packs, with colour, weight, size and position adjustable inside each.

A recorded human is inconsistent by nature. In a course library that is a liability. In a paid feed it is the asset, because inconsistency is what a viewer reads as a person.

There is a route if nobody will film. The avatar module renders a lip-synced presenter from a photo you supply and a voice-over, at 320 credits per rendered minute. That render then goes through the same batch and comes back cut, covered and captioned like any other take.

Paying for two tools costs less than doing either job badly.

Questions people ask

Can I keep a generated presenter and still get a real cut?
Partly. The avatar module renders a lip-synced presenter from a photo and a voice-over at 320 credits per rendered minute, and that output lives inside the project, so it gets the transcript-directed cut, the captions and the footage. What you cannot do is import a rendered video from another tool, because a flattened file has no word timing left to direct.
Why is three minutes the limit on a source take?
Because ads built from short purposeful takes beat ads carved out of long ones, and the ceiling enforces it. Three minutes is roughly 450 spoken words, which is more than a thirty second ad needs with alternatives inside it. If your source is naturally an hour, you are doing extraction, which is a different category of tool.
What does the avatar module actually do?
It renders a lip-synced presenter from one photo you supply and one voice-over, at 320 credits per rendered minute, with generated voice-over at 70 credits a minute. That render then goes through the same batch and comes back cut, covered and captioned. It is built for an ad workflow rather than for a content library.
What if I need both ads and course videos?
Run both and stop treating it as loyalty. A lesson is information for somebody who already chose to watch. An ad is a fight for the next two seconds with somebody who did not. Paying for two tools costs less than doing either badly.
Does anything about my HeyGen library transfer?
The scripts and the knowledge of what a delivered line has to sound like. The rendered videos do not, because the edit needs word-level timing that a finished file no longer has.
Is it right for every team?
It is built for paid vertical ads from a recorded take. If everything you publish is multi-language, longer than three minutes, or square and horizontal, keep a presenter platform for that work and run this alongside it for the ads.

A presenter render is where most tools stop. Cutroom carries it the rest of the way: cut, cutaways on the right words, captions rendered in, and a 9:16 file you can upload today.

Start with one take300 free credits · no card · cancel anytime