Guide
Transcript-based editing: change the words and the video changes with them
Transcript-based editing turns your audio into text in which every word remembers the second it was spoken. Delete a sentence in the text and that stretch of video goes with it. That is the whole definition. The rest of this page is the data it depends on, one paragraph edited three ways, which tools offer it, and where it stops being the right instrument.

- The smallest unit a transcript edit can move
- 1 word
- Moves available on the transcript in Cutroom
- 11
- Vendors publishing a transcript editor, listed below with the date checked
- 4
The definition, and the one piece of data it depends on
Transcribe a recording and you get text. That is a document, and you cannot edit a video with a document.
What makes editing possible is word-level timing. Every word carries the exact moment its audio starts and ends. The word 'expensive' is not a string. It is a string that knows it lives between 00:14.220 and 00:14.910.
Once that mapping exists, an instruction about text becomes an instruction about time. Nobody types a timecode. Remove this sentence means remove that stretch of audio. Cover this phrase means cover those milliseconds.
That is the whole idea. Everything else is what a given tool chooses to do with it.
It feels different to a timeline because your instructions carry no time value at all. On a timeline you say put this clip at 14.2 seconds and hold it for 3.1 seconds. Here you say cover this phrase, and the phrase already knows when it is.
So deleting a line earlier in the take does not break anything after it. Nothing downstream was pinned to a number that has now moved.
Which tools do this, and what each of them calls it
This is not one vendor's feature. Four published a transcript or text-based editor on their own product pages when we checked on 1 September 2026, and they split into two groups.
The larger group is subtractive. You get an accurate transcript and you take material out of it: filler words, a bad sentence, a long pause. That is useful and it is where most tools stop.
What a subtractive pass leaves you with is a trimmed take. A trimmed take is not an ad. It still needs footage placed on the right words, captions styled and timed, a pace decided and an export in the right shape.
The smaller group builds the cut as well as removing from it. That is the distinction worth shopping on, and it is not visible from a feature list that says 'text-based editing' on both sides.
One warning that applies to every tool here: the transcript knows what was said and when, and nothing about what is in the picture. No transcript editor can choose between two takes that say the same words.
- Descript documents editing a recording like a document, in its own help centre. Checked 1 September 2026.
- Adobe publishes text-based editing inside Premiere Pro. Checked 1 September 2026.
- VEED publishes a transcript-based video editing tool page. Checked 1 September 2026.
- CapCut publishes a transcript editing tool page. Checked 1 September 2026.
- Cutroom uses the transcript as the only interface, and builds the cut rather than only trimming it.
Subtractive transcript editing and a drafted cut
Both start from the same word-level timing. They differ in what comes back when you stop typing.
| Most transcript editors | Cutroom | |
|---|---|---|
| Remove a sentence | Yes | Yes |
| Remove filler words | Yes | Yes |
| Place b-roll on a phrase | You do it on a timeline | Highlight the phrase |
| Captions styled and timed | You set them | Six packs, burnt in |
| What you get back | A trimmed take | A finished 9:16 MP4 |
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
One paragraph, edited three ways
Take a real sentence from a founder recording an ad for a bookkeeping app. The transcript reads: 'We were paying an accountant four hundred a month to do something that takes the software about nine seconds, and I found that out in October.'
Edit one. Highlight the phrase 'four hundred a month'. A clip lands over exactly those words and only those words. Not near them. Not at the start of the sentence. Over the 1.4 seconds where he says the number.
Edit two. Delete the clause 'and I found that out in October'. It slows the close. The audio goes and the cut rebuilds around the gap, so the sentence now ends on 'nine seconds'.
Edit three. Mark the word 'nine' for emphasis so it lands larger in the captions. Then change the pace from chill to fast. The whole read re-cuts with shorter gaps, tighter joins and more room for coverage.
Three changes, and you never said a number to the software. You pointed at words.
From audio to a change you can make
Everything hangs on step two. Without word-level timing the transcript is a document and the edit is still manual.
Record the take
One pass, up to three minutes
Transcribe with timing
Every word knows its second
Point at words
Highlight, delete, emphasise
Recompile
The cut rebuilds around the change
Export
20 CR per output minute
Word level is the floor, and that is the honest limit
The resolution is a word. The smallest change you can ask for is a word long, so a cut three frames earlier than the start of a syllable has nowhere to go here. A timeline is the right tool for that and always will be.
In practice it matters less than it sounds. Almost every decision in an ad is already word shaped: where the hook ends, which sentence goes, what gets covered, when the pace changes.
It matters for a small set of jobs. Comedic timing that depends on a specific frame. A music-led edit where the picture hits a beat. A sound cue landing exactly on an action.
Both limits are the price of the speed, and it is a good trade. Word level is coarse enough for the machine to own every time calculation, and owning every time calculation is what lets it rebuild a whole cut in seconds.
If your changes sound like 'cover this sentence with something', this is your instrument. If they sound like 'move that four frames left', it is not.
- The smallest unit is a word, not a frame.
- Instructions carry no time value, so an earlier change never breaks a later one.
- The transcript knows what was said, never what is in the shot.
- There is no timeline, no keyframes and no layers to drop back into.
- A mis-heard word is fixed in the transcript and everything downstream follows.
One take of three minutes in, a 9:16 MP4 out
Here the batch does the building. One take of up to three minutes, and the cut comes back already made: transcribed, trimmed, covered with footage on the phrases that earn it, captioned from one of six packs, paced.
Record one continuous pass on a phone at eye height near a window. Do not stop when you stumble and do not shoot fragments to assemble later. The system wants one unbroken transcript.
The draft that comes back is roughly right and specifically wrong, which is the correct state for a first draft.
Make three changes rather than thirty. The three with the largest effect are the opening line, one badly placed cutaway and the pace. Each change lands in a couple of seconds rather than a couple of minutes.
Cutaway coverage is capped by the compiler at 40 percent of the runtime at chill pace, 45 at normal and 52 at fast, so the speaker stays the centre of the ad however many phrases you highlight.
A batch is 100 credits and export is 20 credits per output minute, so a thirty-second ad from a fresh take is 110 credits. The first three videos are free, over 7 days, with no card, and free exports carry a Cutroom mark across the middle of the picture.
One thirty second ad, in credits
A batch and an export. The trial covers about eight takes and eight exports of that length.
- Export, thirty seconds10 CR
- Batch, one take100 CR
- Free credits in the trial600 CR
Questions people ask
- What is transcript-based editing, in one sentence?
- It is editing video by editing its transcript, which works because every word in that transcript carries the timestamp of the audio it came from. Change the text and the software changes the corresponding stretch of video.
- Which tools offer it?
- On 1 September 2026, Descript, Adobe Premiere Pro, VEED and CapCut each published a transcript or text-based editing feature on their own pages, and Cutroom uses the transcript as its only interface. Most of them let you remove material; fewer build the cut for you.
- What happens when the transcript mishears a word?
- You correct it in the transcript and everything downstream follows, because the captions and the cut both read from the same text. Product names and unusual spellings are the usual offenders, and fixing them once per project is a normal part of the loop.
- Can I edit a video I did not record here?
- It needs the original take rather than a finished file. A rendered video is flattened, with the word-level timing gone, so there is nothing left to point at. Bring the raw recording and it works normally.
- Is it right for every edit?
- No. It is built for spoken video, and it is at its best on vertical ads from one take. Frame-level control, music-led edits and output that is not vertical belong on a timeline.
Point at a word and the video changes. Most transcript tools stop at removing material. Cutroom builds the whole cut and hands back a 9:16 file.