Comparison
HeyGen finishes at the render. An ad starts about there.
A presenter generator turns a script into a person on screen. HeyGen does that in a stack of languages and at company scale. Then the render lands and the work restarts. Here is where each tool stops, what the second stage costs, and how Cutroom returns a finished ad.

- per minute of avatar render in Cutroom
- 320 CR
- per minute of generated voice-over
- 70 CR
- caption packs shipped, no design work
- 6
HeyGen is built for presenter video across a company. Cutroom is built for the ad.
Record once and publish in a stack of languages, with lip movement matched to the new audio. For a company selling in nine markets that is a serious capability.
The API matters too. When video generation has to sit inside a product or a support flow, that is a build. HeyGen is built for the build.
Both are real reasons to buy there, and they are worth naming before anything else.
Neither is what an ad account needs on a Tuesday. An ad account needs finished creatives, in volume, this week.
That is the job Cutroom does end to end. One take goes in, up to three minutes. A 9:16 MP4 comes back with captions burned in and b-roll already placed.
The batch that does it is 100 credits, whatever the take is about. Export is 20 credits per finished minute. Nothing is priced per generated video, so the fourth attempt does not cost what the first one did.
The avatar module is deliberately smaller: one photo plus a voice-over, lip-synced, at 320 credits a minute. One presenter you own, then cut like any other take.
A presenter talking for sixty seconds is a raw ingredient, not an advert
The render arrives and the work restarts. It needs captions, because most of the feed is watched with the sound off.
It needs footage over the claims. Sixty seconds of one framing is where a viewer leaves.
It needs a first line that interrupts. It needs a pace that does not sag at second nineteen.
That second stage is the part that eats the afternoon. A presenter generator does not claim to do it.
So the real comparison is not presenter against presenter. It is a presenter clip against a finished file.
If you are already exporting renders into a second editor to caption and cut them, that second app is what this replaces.
Where the hour after the render actually goes
Cutroom collapses the second stage into marking a transcript. The batch that does it is 100 credits.
- Captioning by hand25 min
- Finding and placing b-roll30 min
- Pacing and trims20 min
- Marking a transcript instead4 min
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Delete the line that sagged and the whole cut rebuilds around the gap
Cutroom transcribes your take and drafts the cut. Back comes a 9:16 MP4 with captions and b-roll already in place.
You then direct it on the words. Highlight a phrase and a clip lands over exactly those words, not near them.
Delete a line and every downstream boundary recomputes. Change the pace from normal to fast. The coverage ceiling moves from 45 to 52 percent with it.
Emphasise a single word so it pops in the captions. Fix a mis-heard word and the burned-in text follows the correction.
Switch the caption pack, colour, weight, size and position from a menu. Six packs ship. None of them needs a designer.
There is no timeline, no keyframes and no layers. None of those is a decision about whether the ad works.
Trim the opening or the ending without touching a track. Write the headline that carries the first three seconds, or take the drafted one.
- Five ways a clip can enter: cut, whip, punch, glitch or sweep
- Music on or off, track chosen, level set under the read
- Coverage ceilings of 40, 45 and 52 percent so the cut stays a person talking
- Export at 20 credits per finished minute, from a take already batched
Three minutes in, 9:16 out, and no timeline waiting underneath
Source over three minutes is refused. This is built for ads, and a three-minute ad is already long.
There is one output shape and it is vertical. That is the shape TikTok, Reels, Shorts and Meta Stories all run.
There is no frame-level control. The compiler owns every boundary in the file, which is why the edit takes four minutes rather than an evening.
Nothing of yours is stamped on the export. No logo, no closing frame.
Either you film, or you own a photo and a voice-over the avatar module can drive.
Those are the facts you plan around on day one. They buy you the thing no presenter tool ships: an ad that is finished when the render is.
Presenter video across a company, or thirty ads a month from one face
Many kinds of video, from many people, in many languages, wired into systems. That is a platform purchase.
A stream of vertical ads from footage that already exists, finished without opening an editor. That is this one.
The arithmetic behind it: winners are 5 to 8 percent of creatives, per Motion's analysis of 550,000+ Meta ads.
Six creatives a month draws about half a winner. Thirty draws two. Nothing about the craft changes that. Only the count does.
A second cut of an existing take costs 20 credits a minute. That price is why the idea you were unsure about gets made anyway.
The trial is seven days and 300 credits, no card. Roughly two finished one-minute ads from fresh takes.
Ten fresh one-minute ads is about 1,200 credits. Basic is 39.99 dollars for 2,500 credits a month, with top-ups at 15 dollars for 1,000.
HeyGen and Cutroom, row by row
Two rows go to HeyGen and they matter to a company rolling out video. The other five are the distance between a render and an ad.
| HeyGen | Cutroom | |
|---|---|---|
| Nine languages from one record | Translation and lip match | One language, one take |
| API and system integration | Built for developers | A browser app, nothing else |
| A file you can upload as it is | A presenter clip, then your edit | 9:16 MP4, ready to run |
| Captions burned in by default | Available, not the core job | Six packs, every export |
| B-roll tied to a phrase | Presenter is the output | Highlight, clip lands there |
| Recut without regenerating | New render, new script | 20 CR a minute, same batch |
| Music sitting under the read | Bring your own track | Generated, 20 CR |
Questions people ask
- Can I cut a HeyGen render inside Cutroom?
- Yes, if it is under three minutes. It goes in as a video take like any other. It gets transcribed, drafted and captioned, and you direct the cut on the transcript. That is the most common way the two tools sit together.
- Does Cutroom do translation or dubbing?
- No. One take, one language, one finished vertical file. A campaign that has to exist in nine markets from a single recording wants a translation-first tool for that job.
- What does the avatar module actually take as input?
- A photo and a voice-over. The result is a lip-synced presenter you can then cut like any other take. Avatar render is 320 credits a minute, generated voice-over is 70 credits a minute, and both sit on top of the 100-credit batch.
- Is it right for every video?
- It is built for vertical ads. Training, onboarding and internal comms want a platform tool. Bring the ads here.
HeyGen hands you a presenter. Cutroom hands you the ad: captions burned in, b-roll on the phrases you marked, a 9:16 file you can upload.