Skip to content

Feature

The transcript is not a preview of the edit. It is the edit

Text based video editing means the words are the control surface, not a shortcut to the trim tool. Delete a line and the cut rebuilds around the gap. Highlight a phrase and footage lands on exactly those words. This page explains how that works, where the idea came from, and the three jobs a timeline still does better.

Surface the whole edit happens on
1
One batch: transcript, cut, b-roll search
100 CR
Longest spoken take accepted
3 min

What the words control: the cut, the footage, the captions, the pace

Upload one spoken take of up to three minutes. It comes back transcribed with word-level timing, already cut, covered with footage and captioned.

From there the words are the controls. Delete a line and the cut rebuilds around the gap. Highlight a phrase and a clip lands over exactly those words, starting on the first and gone on the last.

Change the pace and the whole read is re-cut. Fix a mis-heard word and the captions and the footage search both read the correction.

None of that opens a timeline, because there is no timeline underneath. The transcript is not a friendlier view of the real editor sitting somewhere below. It is the editor.

One instruction, traced through the machine

No step stores a timestamp, which is why an edit above your mark cannot move it.

  1. You mark words

    Highlight, delete, emphasise

  2. Words carry their timing

    From the transcription

  3. The cut recomputes

    From the take, every time

  4. Export

    9:16 MP4, 20 CR a minute

The first generation taught text to delete. Directing needs more verbs

Text based editing arrived as a trim control. A tool transcribes the recording, you delete words, and the matching video goes with them.

Descript made that generation mainstream, and it is a serious product at $16 a month per person on the Hobbyist plan billed annually, per their pricing page, August 2026.

Deleting is one verb, and an edit needs more of them. Where does footage go. What do the captions look like. What happens to the pacing when a paragraph comes out.

In the first generation those answers live on a timeline under the text. The transcript navigates, and the timeline still decides.

The second generation gives the transcript the other verbs. Cover this phrase. Re-cut at this pace. Emphasise this word. When the words hold every instruction, the timeline underneath has nothing left to do, so it is gone.

Two generations of text based editing

The first column improved on scrubbing and deserves the credit for it. The second is a different job description.

Text as a trim controlText as the whole surface
Deleting a sentenceTrims those frames outRebuilds the whole cut
Placing footageBy hand, on the timeline belowHighlight the phrase it covers
CaptionsA separate pass or a pluginTimed from the same transcript
The second versionEdit the last edit againRecomputed from the take
Under the textA timeline you still openNo timeline exists

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

Words are stable addresses. Timestamps are promises that break

Every instruction here points at a word, never at a second. That sounds like a detail and it is the entire mechanism.

A timestamp is only true of one version of the edit. Delete a line and every time value after it is wrong, so a tool built on timestamps has to repair them by guessing.

A word is true of the take itself. Delete a line and the words you marked are exactly where they were, so the cut is recomputed from your instructions rather than repaired.

That is why edit twenty behaves like edit one. Captions do not slide. A clip does not hang over the wrong sentence. There is no accumulated state to corrupt.

It is also why removal works by the line. Pulling a single word out of spoken audio leaves a join you can hear, so cuts land where sentences end and nowhere else.

Three jobs where the timeline is still the right answer

A timeline earns its complexity on three jobs, and pretending otherwise would make the rest of this page worthless.

Frame-accurate work. A beat that has to land on a drum hit, a jump cut placed to the frame. Words are coarser than frames, on purpose.

Several sources in one edit. An interview intercut with a screen recording and archive footage is assembly work, and assembly is what tracks are for.

Masters in other shapes. A landscape YouTube cut or a square version is a timeline export. Everything here composes and exports 9:16 vertical, and nothing else.

The scope on this side is exactly one job. One spoken take of up to three minutes becomes one finished vertical video. If that is your job, the timeline was always too much tool for it.

What a pass costs, and why changing your mind is the cheap part

One batch is 100 credits. It covers the transcription, the directing pass and the footage search across millions of free clips.

Marking the transcript costs nothing. Export is 20 credits per minute of finished video, so a thirty second cut is about 10 credits to download.

Three finished videos are free inside your first seven days, no card, on the 600 credits granted at signup, and they export with a Cutroom mark across the middle. Lite is $19.99 a month for 1,250 credits, nothing on the picture, cancel in one click. That is about eleven short videos a month, made and downloaded.

Basic is $39.99 for 2,500 credits. Premium is $79.99 for 5,000. A $39.99 top-up adds 2,500 when a month runs long.

The number that matters is the second attempt. A different opening line is minutes of marking and 10 credits to export, which is what makes actually trying it worth the bother.

Questions people ask

Is this the same as editing captions?
No. Captions are one output of the transcript. Here the transcript also carries the cut, the footage placement, the pacing and the emphasis. The captions come along because they read from the same word timings.
Can I type new sentences into the transcript?
No. You can fix a word the transcription mis-heard, because the audio proves what you said. You cannot type words you never spoke. The take is the source, and the editor directs it rather than rewriting it.
What if I delete the wrong line?
Put it back. A deletion is an instruction, not surgery on a file. The original take is untouched and the cut is recomputed from your current instructions every time.
Why does it refuse to cut in the middle of a sentence?
Because you would hear it. A word pulled out of continuous speech leaves a seam in the breath and the room tone. Cutting at sentence ends is the difference between edited and chopped.
Who should stay on a timeline?
Anyone assembling several sources, anyone who needs frame-level control, and anyone delivering landscape or square masters. This paradigm fits one spoken take becoming one vertical video, and it does not stretch past that shape.

The edit was always described in words first. This is the version where the description is the edit.

3 videos free, no card3 finished videos free in your first 7 days, no card. They carry a Cutroom mark; Lite at $19.99/month removes it