Cutroom
Features
One spoken take goes in. These pages each explain one thing the engine does to it, including where the honest limits sit.
Upload one take. *A finished 9:16 ad comes back*
Ad editors hand you a timeline and expect you to build the ad on it. This one hands back a finished 9:16 MP4 and asks you to mark words. What follows is what goes in, what comes back, and what each step charges. By the end you can decide whether one phone take is enough to ship this week's test.
Direct the ad by marking words, *not by dragging clips*
You edit the whole ad by marking words on a transcript of your take. There is no timeline, no dragging, and nothing to type. This page lists the eleven edits you can make, why the list is fixed, and what one finished 9:16 ad costs.
Highlight a phrase and *the clip covers those exact words*
The machine searches, picks the shots, and cuts each one to the words it covers, all inside the 100 credit batch. Here is where the footage comes from, how much of the ad it may cover, and what a swap costs.
One photo, one audio track, *and a presenter renders*
One front-facing photo and one audio track return a lip-synced presenter take at 350 credits a rendered minute. That take runs through the editor like filmed footage and exports as a finished 9:16 ad. Here is what makes a photo work, what the audio decides, and what a campaign costs.
Describe the sound in plain words. *Fifty credits a track*
You describe the sound in ordinary language, add lyrics if you want them sung, and a finished track comes back for 50 credits at any length. Here is what to write, how loud to set it under a read, and why four candidates beat one brief.
Word-timed captions, *rendered into the file itself*
Captions here come out of the same transcription that drives the cut, so they land on the beat of your speech rather than a syllable behind it. Here is where the timing comes from, which of the six packs to pick, and why captioning adds nothing to the bill.
One messy take in, *a cut ad back in minutes*
One messy talking-head take goes in and a cut vertical ad comes back in minutes. The throat-clearing and the tangent get fixed in the words rather than on a timeline. Here is what four moves on a transcript do to a rambling read, and what a finished ad costs.
Every export is 9:16, *composed rather than cropped*
Every export here is a 9:16 MP4 composed for the vertical feed, at 20 credits per finished minute. Sound off, small screen, platform buttons on top: the frame is built for all of it. Here is what the file contains, why composing beats cropping, and how to keep the frame safe across apps.
Four lines sung back to you, *fifty credits an attempt*
You write the lyric, describe the sound in ordinary language, and a sung jingle comes back for 50 credits at any length. Here is what to write, how many attempts to run, and where six sung seconds sit in a thirty second ad.
Paste your own lyric. *It comes back sung and mixed*
Paste your lyrics, describe the sound in plain language, and a finished track comes back sung, arranged and mixed for 50 credits. Here is how to write lines that sing cleanly, why you regenerate instead of revising, and what three attempts cost.
Leave the lyrics empty and *the track comes back wordless*
Leave the lyric field empty, describe the sound, and a finished instrumental comes back for 50 credits at any length. This page covers what to write in the description, why a sung vocal fights a spoken read, and how loud to set the music under your ad.
One take in, one finished ad out, *about 110 credits*
One talking-head take of up to three minutes goes in. A finished 9:16 MP4 comes back, cut, captioned, with footage on the words that needed showing and a headline drafted. Here is what you have to film, what each step charges, and how many ads a month a plan buys.
You do the talking. *It does the cutting, for 100 credits*
Film one phone take of up to three minutes, fumbles left in, and a finished vertical UGC ad comes back cut, captioned and covered. The words stay yours, because that is what makes the format work. Here is what to say, how much footage to put over it, and what twenty attempts cost.
Two ways to keep your face out of the video, *and both end in a finished ad*
Two routes here keep your face out of the ad. One renders a presenter from a photo of somebody who agreed in writing, at 350 credits per minute. The other covers as much of the ad as the pace allows, up to 52 percent. Here is what each costs and which one your script needs.
The same face in ad twenty, *from the same single photo*
One front-facing photo of somebody who agreed in writing renders the same presenter in every ad, at 350 credits a minute. A spokesperson works by the fifth exposure, and this removes what usually breaks first, the bookings. Here is what consent has to cover and what a campaign costs.
Every edit points at a word, *so nothing breaks when you change your mind*
Here the take is transcribed with word-level timing and every instruction points at a word. Cover this phrase. Remove this line. Nothing drifts when you change your mind. Here is why that survives the second edit, what transcription gets wrong, and what one take costs to run.
Dead air comes out. *No cut ever lands inside a word*
Cutroom shortens the long silences in one spoken take and, at the fast pace, drops the flagged ums, on its way to turning that take into a finished 9:16 ad. It is not a trimmer you feed arbitrary files. This page gives the exact pause lengths it cuts and keeps, in milliseconds, and why a cut mid-word is the one edit it refuses to make.
The punch-in effect, *aimed at the word that carries the claim*
Punch-in zooms here are placed by the machine, not by hand. One may land at a sentence start or on a word you emphasised, it scales between 1.06 and 1.12, and two can never fire within six seconds of each other. This page explains where they land, when they buy retention, and when they are just noise with momentum.
The ums go, and *the transcript shows you every word that went*
Cutroom's transcriber flags the hesitation sounds in your take, and the fast pace drops them from the cut. You audit the removals by reading rather than by scrubbing, because every dropped word stays visible in the transcript. Here is exactly what gets removed, what deliberately stays, and the one thing this will not do to your sentences.
The hook gets drafted *from the words already in your take*
Cutroom writes the opening headline of your ad from the transcript of your own take, or you write it yourself and it is one field. Both routes need the same raw material, a claim worth putting first. This page gives you the eight hook patterns the director draws on, with a worked example for each, and they work on paper before you ever open a tool.
The transcript is not a preview of the edit. *It is the edit*
Text based video editing means the words are the control surface, not a shortcut to the trim tool. Delete a line and the cut rebuilds around the gap. Highlight a phrase and footage lands on exactly those words. This page explains how that works, where the idea came from, and the three jobs a timeline still does better.
Add music to a video: *the track is the easy half, the mix is the job*
If your video is finished and just needs a track, the editor already on your phone does that at no cost. This page is for the other case: music under a person talking, in a video that is still being built, where the wrong bed eats the words. Here is how a described mood becomes a bed for 50 credits, and how the mix keeps the voice on top.
Automatic is not the same as out of your hands. *Every decision shows on the transcript*
Most automatic editors hand you a rendered file and one lever, regenerate. This one makes every first decision itself, then shows each one on the transcript of your take, where reversing it means pointing at it. Here is what gets decided for you, where you read it, and what a finished pass costs.
No camera and no face, *and the script is still your job*
The numbers first, because faceless is sold as effortless and it is not. A presenter renders from one photo and one audio track at 350 credits a minute, on Basic at $39.99 a month. A generated read is 70 credits a minute. What no module replaces is the script, and this page is honest about that part.
Caption styles: six complete looks with every ad, *switch in one click*
Six video ad caption styles ship with every batch, each a full look rather than a font. This page covers what a pack decides, which category each suits, and what changes when you switch.
Clean UGC captions: *one soft line, four words, no effects*
The clean UGC pack puts one rounded line on screen, four words at a time, cuts and sweeps only, no sound effects. It is built to read as posted rather than produced. This page covers what lands on screen, why the restraint works, and which ads it suits.
Bold impact captions: three words in heavy capitals, *built for price and deadline offers*
The bold impact pack puts three words on screen at a time, in heavy capitals with whip and punch cuts. It is built for offers with a price and a deadline that have to be read in one second. This page covers what lands on screen, why the ceiling is three words, and which ads it wins.
Karaoke captions: the whole line stays readable, *each spoken word lights up*
Karaoke captions show the whole line and light the word being said. The eye reads ahead and never loses its place. That is worth a lot on anything with steps in it. This page covers what lands on screen, why mid-frame is deliberate, and which scripts it suits.
Editorial captions: *five words, serif, and nobody rushing you*
The editorial pack sets serif type in a lower third, up to five words at a time, crossfades and a chill pace. It is built for purchases people think about. This page covers what lands on screen and what five words lets you write.
Fast native captions: *a text strip on top, subtitles below*
The fast native pack opens on a text strip across the top of the frame, the meme caption format viewers read as a post rather than an ad. Captions run below it three words at a time, with whip and glitch transitions. This page covers what lands on screen, what the strip is for, and which ads it holds together.
Minimal lux captions: *small, lowercase, and almost still*
The minimal lux pack puts small lowercase captions with wide letter spacing on screen, almost no zoom and no sound effects. It is the caption look luxury and beauty brands use to make a product read as expensive. This page covers what lands on screen, why the stillness works, and what footage it asks of you.
Sound off captions: every ad is captioned automatically, *here is how to write for muted viewers*
Most of your audience will meet the ad with the sound off. Every pack here captions it automatically inside the 100 credit batch. What changes is the writing, because a muted viewer reads three to five words at a time. This page covers the word ceilings, the opening group, and the test that catches most failures.
AI music for TikTok: *describe it, and a track comes back*
AI music for TikTok exists, and it is one screen. Type how the track should sound, pick a length, and a finished piece comes back for a flat 50 credits, sung or instrumental. This page covers what you type, what it costs, where the terms stand, and how to brief a track that survives a phone speaker.
AI music for Instagram Reels: *describe the sound, get a track for 50 credits*
AI music for Instagram Reels is one step: describe the sound and a finished track comes back for 50 credits, 15 to 120 seconds. This page covers what the module makes, why the in-app library rarely covers paid ads, how the terms stand, and how to brief a track for a phone speaker.
AI music for YouTube: *a finished track for 50 credits, up to 120 seconds*
Describe the sound and a finished track returns for a flat 50 credits, at 15 to 120 seconds. Two minutes is the ceiling, right for an ad and wrong for a long upload. This page covers what it makes, how matching systems see it, and where one track pays for itself for a year.
Background music for ads: *describe the sound, get a finished track for 50 credits*
Background music for an ad takes one step: describe the sound, pick 15 to 120 seconds, and a finished track comes back for 50 credits. This page covers how to write a description that leaves room for the voice, the phone test, and the ads that are better with no music at all.
A royalty free music alternative: *describe the track and it gets made for 50 credits*
The alternative to royalty free music is a track generated for you: describe the sound and a finished piece comes back for 50 credits, 15 to 120 seconds. This page covers what that fixes, where licence tiers catch people, and the two jobs a stock library still does better.
AI jingle for your brand: *write the words, get a sung track for 50 credits*
A jingle costs 50 credits to make. Type the words, describe the sound, pick 15 or 30 seconds, and a sung track comes back. The expensive part is running the same fifteen seconds for a year. This page covers how to make one, what the words have to do, and how to tell whether your brand will keep one.
Custom song from lyrics: *paste your words, get a sung track for 50 credits*
Paste your lyrics, describe the sound, pick a length up to 120 seconds, and a sung track comes back for 50 credits. The words decide whether anybody can follow it, and that is the part most people get wrong. This page covers the word budget, writing for half attention, and when to leave the lyrics box empty instead.
Instrumental for a voice-over: *describe the sound, get a vocal-free track for 50 credits*
An instrumental for a voice-over costs 50 credits: describe the sound, leave the lyrics box empty, and a track with no vocals comes back at 15 to 120 seconds. This page covers how to write the description, why lowering the volume rarely fixes a clash, and how to match the track to the voice.
Writing jingle lyrics: *fifteen words, name in the first three*
A sung line carries about a third of the words a spoken one does, so a fifteen second jingle is about fifteen words. That budget decides the writing. This page covers where the brand name sits, why repetition is the mechanism, and the two free tests to run before you spend 50 credits.
Music under a voice-over works when *it avoids one frequency band*
A speaking voice occupies roughly 500 Hz to 4 kHz, and a bed works when it stays out of that band. Mood, genre and tempo are downstream of that one fact. This page covers why masking is not about volume, what to type to get a bed that survives, and how a phone speaker changes the answer.
Copyright safe music for ads: *three routes, and where each one breaks*
Three routes exist for music behind a paid ad: a platform library, a stock licence, or a track generated for you at 50 credits. Each fails in a different place. This page says where each breaks, what to check before you use it, and the four records that answer a question asked two years later.
How long should ad music be: *as long as the ad, and every length costs the same*
Ad music should run exactly as long as the ad. The lengths on offer are 15, 30, 60, 90 and 120 seconds, all at 50 credits, so generate the length your ad already is. This page covers what each length is for, why a track that ends beats one that stops, and the extra word limit on sung tracks.
AI music for podcasts: *a theme, a sting and an outro at 50 credits each*
A show needs three tracks: a theme, a sting and an outro. Describe the sound, pick a length, and each comes back for 50 credits, so 150 credits covers a show's whole sonic identity. This page covers how to brief a theme you can live with, and why a bed under a conversation costs you words.
AI music for product videos: *describe the sound, get a track for 50 credits*
Describe the sound and a finished track returns for 50 credits, at 15 to 120 seconds. Before you generate one, listen to the product. A click or a pour is evidence, and evidence beats production values. This page covers that test, what to type when the product is silent, and how one description covers forty ads.
Instrumental or vocals: *a listener follows one set of words*
Use a sung track when nobody speaks in the ad, and an instrumental under any voice-over. A listener cannot follow two sets of words at once, and both options cost 50 credits. This page covers when each one wins, the one exception that works, and what the two-track pattern costs.
When music fights your voice-over: *six symptoms, six fixes*
The ad sounds muddy and the instinct is to pull the music down. Six different faults produce that feeling. Only one of them is the level. This page names all six, gives the fix for each, and prices it. By the end you will be able to diagnose a muddy ad in about a minute, without learning to mix.
AI presenter video that arrives as *a finished ad*
Presenter tools solved one problem. A photo and a voice file now make a convincing video of someone talking. They did not solve the next part. One unbroken shot of a face is not an advert, and most people find that out after paying. Here is how the format works, what one clip costs, and how Cutroom finishes the job.
Photo to video avatar: *seven rules the portrait has to pass*
One photo and one audio file come back as a lip-synced video of that person talking, at 350 credits a minute. Most renders that fail, fail on the photo. Here are the seven rules the portrait has to pass, and how Cutroom turns the render into a finished ad.
Talking head generator: *what goes in, what comes back, what it costs*
Talking head generators all do the same first thing well. A portrait and an audio file return a video of that person speaking. Almost none of them do the second thing. One unbroken shot of a face is not an ad. Here is what the render needs, what a minute costs, and how Cutroom finishes the file.
AI UGC avatar ads: the avatar carries the words, *real footage shows the product*
An AI UGC avatar can speak your script, but it has no hands, so it cannot show your product. The ads that convert put real footage under the rendered voice. Here is what the render gives you, how one batch adds that footage, and what twenty finished attempts cost.
Faceless brand videos: *two ways to a finished ad without your face on camera*
Plenty of brands sell well with no founder on screen. The block is rarely the idea. It is that every editor still expects a face to cut around. Two routes here do not. Here is the filmed route at 100 credits a batch, the rendered route at 350 credits a minute, and which one your script wants.
AI spokesperson for ecommerce: *one face, forty listings*
Forty SKUs means forty scripts. Nobody stands in front of a camera forty times. A rendered spokesperson removes that constraint at 350 credits a finished minute. Here is what one SKU ad costs, how a single render fronts many listings, and where to spend the credits first.
Lip sync video generator: *your audio decides how good the sync looks*
A lip sync video generator turns one photo and one audio file into a video of that person speaking. The sync quality comes from your recording, not from the model. Here is what the render takes, what a minute costs, the four recordings that break the sync, and how the clip becomes a finished ad.
AI video without filming: *write a script, get a finished ad, no camera needed*
You can make a finished video ad without filming anything. A portrait and a generated voice become a talking presenter, and the editor adds captions and b-roll, all from one desk in 9:16. Here is the route priced step by step, what the same ad costs filmed, and where no camera wins outright.
How AI avatars work: *the photo is the least important input*
An avatar render has three inputs. People rank them in the wrong order. The audio drives the mouth. The script decides whether anybody believes it. The portrait mostly rules options out. Here is the pipeline priced step by step, what each input controls, and the step after the render that turns a take into an ad.
When an avatar beats filming: *five cases and a sorting test*
Both routes make the same finished 9:16 ad and both live in one account. A render is about 393 credits for forty seconds. Filming the same ad is about 114. Here are the five cases where the render is worth that difference, a ten minute test that sorts your script list, and where the rest of your scripts should go.
Film it yourself: *the six lines an AI avatar cannot deliver*
If somebody is willing to go on camera, filming usually beats the avatar. One phone take of up to three minutes batches for 100 credits and comes back cut, with captions and b-roll on it, and the export adds 20 credits a minute. Here is what the filmed route costs, the six sentences a rendered face cannot carry, and when to render instead.
Writing for an avatar: *nine rules and one read aloud test*
Half of what people call a bad avatar is a good render of writing meant for the eye. The render did what it was told. The script told it the wrong thing. Here are the nine rules a script has to pass, the free test that catches almost every failure, and the length the credits and the attention both agree on.
Avatar video for product demos: *the avatar explains, real footage shows the product*
An avatar cannot hold your product, because a rendered frame has no hands in it. A demo ad still works: the avatar explains while real footage shows the product. Here is what the render gives you, where the product footage comes from, and how one batch puts it on the exact words that need it.
AI avatars for courses: *render the ad, film the lesson*
Course creators arrive here wanting rendered lessons. The arithmetic points them somewhere better. A render bills 350 credits per finished minute and caps at five minutes of audio. Here is the lesson maths, the four places a presenter pays for itself, and the one asset worth spending the credits on.
Why avatar ads feel fake: *only one of six causes is the model*
People switch avatar tools to fix this and the ads still feel fake. The model is the sixth cause out of six. Ranked by damage, the list runs script, length, unbroken frame, audio, portrait, model. Here is each cause, the fix, and what the fix costs. Four of the six are free and the fifth is one batch.
Voice-over for avatar video: *three routes, one of them free*
The audio decides how a render looks. Three routes feed it. Generate a read at 70 credits a minute, record your own for nothing, or hire somebody. Here is what each route does to the sync, what each costs against the render it feeds, and how to pick one in under a minute.
5 Reasons Creators Let Cutroom Do Their *B-Roll Edits*
You upload one talking video. The b-roll lands on the words that need it, and the captions and the cuts come with it.
One take in. A finished ad out. See it on your own footage before you decide anything.