Guide

Transcript-based editing means you edit the words, not the clips

Transcript-based editing means the software turns your audio into text and remembers which second every word came from. You change the video by changing the text. Delete a line and that audio goes with it. Highlight a phrase and footage lands over exactly those words. This page covers the timing data underneath, one paragraph edited three ways, and the point where a text editor becomes a finished cut. By the end you will know whether pointing at words suits how you think.

Keyframes, layers and tracks involved
0
Moves available on the transcript
11
Longest take the editor accepts
3 min

The definition, and the one piece of data it depends on

Transcribe a recording and you get text. That is a document, and you cannot edit a video with a document.

What makes editing possible is word-level timing. Every word in the transcript carries the exact moment its audio starts and ends. The word 'expensive' is not a string. It is a string that knows it lives between 00:14.220 and 00:14.910.

Once that mapping exists, an instruction about text becomes an instruction about time. Nobody types a timecode. Remove this sentence means remove that stretch of audio. Cover this phrase means cover those milliseconds.

That is the whole idea. Everything else is what a given tool chooses to do with it.

It feels different to a timeline because your instructions carry no time value at all. On a timeline you say put this clip at 14.2 seconds and hold it for 3.1 seconds. Here you say cover this phrase, and the phrase already knows when it is.

So deleting a line earlier in the take does not break anything after it. Nothing downstream was pinned to a number that has now moved.

One paragraph, edited three ways

Take a real sentence from a founder recording an ad for a bookkeeping app. The transcript reads: 'We were paying an accountant four hundred a month to do something that takes the software about nine seconds, and I found that out in October.'

Edit one. Highlight the phrase 'four hundred a month'. A clip lands over exactly those words and only those words. Not near them. Not at the start of the sentence. Over the 1.4 seconds where he says the number.

Edit two. Delete the clause 'and I found that out in October'. It slows the close. The audio goes and the cut rebuilds around the gap. The sentence now ends on 'nine seconds', which is a stronger place to end it.

Edit three. Mark the word 'nine' for emphasis so it lands larger in the captions. Then change the pace from chill to fast. The whole read re-cuts with shorter gaps, tighter joins and more room for coverage.

Three changes, and you never said a number to the software. You pointed at words.

That is what people mean when they say they stopped feeling like they were operating software. The decisions are about the ad now, rather than about the tool.

From audio to a change you can make

Everything hangs on step two. Without word-level timing the transcript is a document and the edit is still manual.

  1. Record the take

    One pass, up to three minutes

  2. Transcribe with timing

    Every word knows its second

  3. Point at words

    Highlight, delete, emphasise

  4. Recompile

    The cut rebuilds around the change

  5. Export

    20 CR per output minute

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

Word level is the floor, and it is why a rebuild takes seconds

The resolution is a word. The smallest change you can ask for is a word long.

A cut three frames earlier than the start of a syllable has nowhere to go here. A timeline is the right tool for that and always will be.

In practice it matters less than it sounds. Almost every decision in an ad is already word shaped. Where the hook ends. Which sentence goes. What gets covered. When the pace changes. All of those land on word boundaries anyway.

It matters for a small set of jobs. Comedic timing that depends on a specific frame. A music-led edit where the picture hits a beat. A sound cue landing exactly on an action.

The second fact is what the transcript describes. It knows what was said and when. It knows nothing about what is in the picture, so it cannot choose between two takes that say the same words.

Both of those are the price of the speed, and it is a good trade. Word level is coarse enough for the machine to own every time calculation. Owning every time calculation is what lets it rebuild a whole cut in seconds.

  • The smallest unit is a word, not a frame.
  • Instructions carry no time value, so an earlier change never breaks a later one.
  • The transcript knows what was said, never what is in the shot.
  • There is no timeline, no keyframes and no layers to drop back into.
  • A mis-heard word is fixed in the transcript and everything downstream follows.

Most transcript tools remove. This one builds the cut as well.

The common version of transcript editing is subtractive. It gives you an accurate transcript and lets you take material out. Filler words, a bad sentence, a long pause. That is useful, and it is where most tools stop.

What it leaves you with is a trimmed take. A trimmed take is not an ad. It still needs footage placed on the right words, captions styled and timed, a pace decided and an export in the right shape.

Here the batch does the building. One take of up to three minutes, 100 credits, and the cut comes back already made. Transcribed, trimmed, covered with footage on the phrases that earn it, captioned from one of six packs, paced.

Then the transcript stays the interface for changing it. Eleven moves live there, and none of them is typing into a free-text field.

Cutaway coverage is capped by the compiler at 40 percent of the runtime at chill pace, 45 at normal and 52 at fast. The speaker stays the centre of the ad however many phrases you highlight.

The output is a 9:16 MP4 at 20 credits per output minute. One shape, made well.

  • Highlight a phrase and a clip lands over exactly those words.
  • Swap the clip in any spot, or search millions of free ones.
  • Delete a line you do not want and the cut rebuilds around it.
  • Change the pace and the machine re-cuts the whole thing.
  • Change how each clip enters: Cut, Whip, Punch, Glitch or Sweep.

One take of three minutes in, a 9:16 MP4 out

The loop is short enough to describe in a paragraph, which is the point.

Record one continuous pass on a phone, up to three minutes, at eye height near a window. Do not stop when you stumble. Do not shoot fragments to assemble later. The system wants one unbroken transcript.

Upload and run a batch for 100 credits. The draft that comes back is roughly right and specifically wrong, which is the correct state for a first draft.

Make three changes rather than thirty. The three with the largest effect are the opening line, one badly placed cutaway and the pace. Highlight, delete, re-pace. Each change lands in a couple of seconds rather than a couple of minutes.

Export at 20 credits per output minute. A thirty second ad is 10, so one ad from a fresh take is 110 credits.

The trial is 7 days and 300 credits with no card. That covers two takes and four finished thirty second ads, which is enough to find out whether pointing at words suits how you think.

One thirty second ad, in credits

A batch and an export. The trial covers two takes and four exports of that length before a card is involved.

  • Export, thirty seconds10 CR
  • Batch, one take100 CR
  • Free credits in the trial300 CR

Questions people ask

What happens when the transcript mishears a word?
You correct it in the transcript and everything downstream follows, because the captions and the cut both read from the same text. Product names and unusual spellings are the usual offenders, and fixing them once per project is a normal part of the loop.
Can I edit a video I did not record here?
It needs the original take rather than a finished file. A rendered video is flattened, with the word-level timing gone, so there is nothing left to point at. Bring the raw recording and it works normally.
Is this the same as removing filler words automatically?
That is one subtractive feature built on the same data. Transcript-based editing is the broader idea: any instruction about words becomes an instruction about time. Placing footage on a phrase and re-pacing a whole cut use the same mapping as removing an 'um'.
How long does one change take to see?
Seconds, because nothing has to be recalculated by hand. The machine owns every time value in the project, so deleting a line in the middle rebuilds the whole cut around the gap without anything after it moving out of place.
Is it right for every edit?
It is built for vertical ads from a spoken take. Frame-level control, music-led edits and output that is not vertical belong on a timeline. The tell is what your changes sound like. Cover this sentence with something is a transcript sentence.

Point at a word and the video changes. Most transcript tools stop at removing material. Cutroom builds the whole cut and hands back a 9:16 file.

Start with one take300 free credits · no card · cancel anytime