Feature
The transcript is not a preview of the edit. It is the edit
Text based video editing means the words are the control surface, not a shortcut to the trim tool. Delete a line and the cut rebuilds around the gap. Highlight a phrase and footage lands on exactly those words. This page explains how that works, where the idea came from, and the three jobs a timeline still does better.
- Surface the whole edit happens on
- 1
- One batch: transcript, cut, b-roll search
- 100 CR
- Longest spoken take accepted
- 3 min
What the words control: the cut, the footage, the captions, the pace
Upload one spoken take of up to three minutes. It comes back transcribed with word-level timing, already cut, covered with footage and captioned.
From there the words are the controls. Delete a line and the cut rebuilds around the gap. Highlight a phrase and a clip lands over exactly those words, starting on the first and gone on the last.
Change the pace and the whole read is re-cut. Fix a mis-heard word and the captions and the footage search both read the correction.
None of that opens a timeline, because there is no timeline underneath. The transcript is not a friendlier view of the real editor sitting somewhere below. It is the editor.
One instruction, traced through the machine
No step stores a timestamp, which is why an edit above your mark cannot move it.
You mark words
Highlight, delete, emphasise
Words carry their timing
From the transcription
The cut recomputes
From the take, every time
Export
9:16 MP4, 20 CR a minute
The first generation taught text to delete. Directing needs more verbs
Text based editing arrived as a trim control. A tool transcribes the recording, you delete words, and the matching video goes with them.
Descript made that generation mainstream, and it is a serious product at $16 a month per person on the Hobbyist plan billed annually, per their pricing page, August 2026.
Deleting is one verb, and an edit needs more of them. Where does footage go. What do the captions look like. What happens to the pacing when a paragraph comes out.
In the first generation those answers live on a timeline under the text. The transcript navigates, and the timeline still decides.
The second generation gives the transcript the other verbs. Cover this phrase. Re-cut at this pace. Emphasise this word. When the words hold every instruction, the timeline underneath has nothing left to do, so it is gone.
Two generations of text based editing
The first column improved on scrubbing and deserves the credit for it. The second is a different job description.
| Text as a trim control | Text as the whole surface | |
|---|---|---|
| Deleting a sentence | Trims those frames out | Rebuilds the whole cut |
| Placing footage | By hand, on the timeline below | Highlight the phrase it covers |
| Captions | A separate pass or a plugin | Timed from the same transcript |
| The second version | Edit the last edit again | Recomputed from the take |
| Under the text | A timeline you still open | No timeline exists |
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
Words are stable addresses. Timestamps are promises that break
Every instruction here points at a word, never at a second. That sounds like a detail and it is the entire mechanism.
A timestamp is only true of one version of the edit. Delete a line and every time value after it is wrong, so a tool built on timestamps has to repair them by guessing.
A word is true of the take itself. Delete a line and the words you marked are exactly where they were, so the cut is recomputed from your instructions rather than repaired.
That is why edit twenty behaves like edit one. Captions do not slide. A clip does not hang over the wrong sentence. There is no accumulated state to corrupt.
It is also why removal works by the line. Pulling a single word out of spoken audio leaves a join you can hear, so cuts land where sentences end and nowhere else.
Three jobs where the timeline is still the right answer
A timeline earns its complexity on three jobs, and pretending otherwise would make the rest of this page worthless.
Frame-accurate work. A beat that has to land on a drum hit, a jump cut placed to the frame. Words are coarser than frames, on purpose.
Several sources in one edit. An interview intercut with a screen recording and archive footage is assembly work, and assembly is what tracks are for.
Masters in other shapes. A landscape YouTube cut or a square version is a timeline export. Everything here composes and exports 9:16 vertical, and nothing else.
The scope on this side is exactly one job. One spoken take of up to three minutes becomes one finished vertical video. If that is your job, the timeline was always too much tool for it.
What a pass costs, and why changing your mind is the cheap part
One batch is 100 credits. It covers the transcription, the directing pass and the footage search across millions of free clips.
Marking the transcript costs nothing. Export is 20 credits per minute of finished video, so a thirty second cut is about 10 credits to download.
Three finished videos are free inside your first seven days, no card, on the 600 credits granted at signup, and they export with a Cutroom mark across the middle. Lite is $19.99 a month for 1,250 credits, nothing on the picture, cancel in one click. That is about eleven short videos a month, made and downloaded.
Basic is $39.99 for 2,500 credits. Premium is $79.99 for 5,000. A $39.99 top-up adds 2,500 when a month runs long.
The number that matters is the second attempt. A different opening line is minutes of marking and 10 credits to export, which is what makes actually trying it worth the bother.
Questions people ask
- Is this the same as editing captions?
- No. Captions are one output of the transcript. Here the transcript also carries the cut, the footage placement, the pacing and the emphasis. The captions come along because they read from the same word timings.
- Can I type new sentences into the transcript?
- No. You can fix a word the transcription mis-heard, because the audio proves what you said. You cannot type words you never spoke. The take is the source, and the editor directs it rather than rewriting it.
- What if I delete the wrong line?
- Put it back. A deletion is an instruction, not surgery on a file. The original take is untouched and the cut is recomputed from your current instructions every time.
- Why does it refuse to cut in the middle of a sentence?
- Because you would hear it. A word pulled out of continuous speech leaves a seam in the breath and the room tone. Cutting at sentence ends is the difference between edited and chopped.
- Who should stay on a timeline?
- Anyone assembling several sources, anyone who needs frame-level control, and anyone delivering landscape or square masters. This paradigm fits one spoken take becoming one vertical video, and it does not stretch past that shape.
The edit was always described in words first. This is the version where the description is the edit.