Guide
A transcript edit works on the words. The time math happens out of sight
A transcript edit is a cut made by working on the text of what was said instead of on a timeline. Delete a sentence in the transcript and the matching stretch of video is removed. The words are the interface, and the frame arithmetic happens out of sight. This page covers why that works for spoken video and where it stops working.
- The precision the underlying timing is stored at
- 1 word
- Tracks, layers or keyframes involved
- 0
- Things it needs: speech, and an accurate transcript
- 2
Spoken video is structured by sentences, so edit the sentences
A talking video's real structure is not visual. It is the argument: this sentence, then that one, minus the tangent in the middle.
A timeline shows none of that. It shows a waveform, and finding the tangent means scrubbing back and forth listening for it.
A transcript shows exactly that. The tangent is three lines of text, and removing it is the same gesture as deleting them from an email.
What one deletion actually does
The editor's gesture is textual. Everything after it is mechanical and repeatable.
Words transcribed
Timing stored per word
You delete a line
A text gesture
The cut rebuilds
Around the removal
Preview updates
Same result every time
What it is good at, and what it is not
It is built for anything carried by speech. Ads, explainers, updates, testimonials: content where the words are the video.
It is the wrong tool for frame-level work. A montage cut to a beat, a colour grade, an action sequence: those decisions live below the word, and a transcript cannot see them.
The honest test is one question. If the video would still make sense as an audio file, transcript editing fits it. If the pictures are the point, use a timeline.
Speed is the other reason people switch. Reading a page of text takes a minute; scrubbing the same take on a timeline takes ten, because you cannot skim audio. The transcript turns a listening job into a reading job, and reading is the faster skill.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Under the hood, an edit is an operation, not a hand-move
The transcription stores a start and end time for every word. A text gesture like deleting a line translates into an instruction against those times.
In Cutroom that instruction set is deliberately closed. Delete a line, change the pace, swap a clip, fix a mis-heard word: each is a defined move, and the whole cut recompiles after every one.
The practical consequence is repeatability. The same take with the same edits produces the same video, which is not something hand-dragged timelines can promise.
Questions people ask
- How accurate is a transcript edit?
- Accurate to the word, because the timing is stored per word. What it does not offer is frame-level trimming inside a word, which is timeline territory and rarely needed for spoken content.
- What if the transcript mishears a word?
- You correct the word and the timing underneath it stays put. A mis-heard word is a caption problem, not a cut problem, so fixing the text fixes what viewers read.
- Can I edit a video with no speech this way?
- No. With nothing said there is nothing to transcribe, and the interface is the transcript. Music-led or visual-first edits belong in a conventional editor.
Point at the words, and let the machine do the time arithmetic. That trade is the entire product here, and it only makes sense for video that talks.