Guide

The transcript is the timeline. The words carry the timecode

You edit the transcript instead of the track. Every word carries the timing of the audio it came from. Delete a line and that audio goes, and the picture rebuilds around the gap. Highlight a phrase and footage lands over exactly those words. This page covers the timing data that makes it work, the trade it asks for, and the four jobs it hands back to a timeline. By the end you will know whether to point at words or at clips.

Camera and monitor setup for video production and photography editing in a studio.
Photo by Jakub Zerdzicki on Pexels
Keyframes, layers and tracks involved
0
Maximum levers, by product law
7
Transitions: Cut, Whip, Punch, Glitch, Sweep
5

Word-level timing is what makes the text an index into the video

Transcription produces more than text. It produces a start and end time for every individual word. That turns the transcript into an index into the recording.

Once that index exists, a text operation becomes a video operation. Selecting words selects a region of the take. Deleting a sentence deletes the audio it names.

The compiler owns all the time arithmetic from there. It knows where each word sits, so it rebuilds the sequence whenever the text changes. It never asks you for a frame number.

That is the mechanism in full. It explains the speed and it explains the resolution. Word level is the finest available, so a cut three frames earlier is not something you ask for.

It also explains why the source has to be a continuous spoken take. With no speech there is no index. With no index there is nothing to edit against.

The index has one more useful property. Words are the addresses, so the description of an edit stays readable by a person. That is rare in video and valuable when something goes wrong.

You see what the system decided in the same terms you gave it. No list of frame numbers. No guessing which one was the mistake.

How a text action becomes a video action

Every step after the first is arithmetic. You never touch a frame number, which is what makes a rebuild take seconds.

  1. Transcribe

    Timing per word

  2. Act on the words

    Delete, highlight, re-pace

  3. Recompile

    Machine owns the time math

  4. Export

    9:16 MP4, 20 CR per output minute

Directing decisions in, editing decisions out, and that is the whole trade

A timeline lets you express anything, at the cost of expressing everything. Every cut, every cutaway placement, every caption timing nudge passes through your hands.

A transcript editor accepts a smaller vocabulary and the machine does the assembly. You say cover this sentence. You do not say place this clip at this time for this duration.

The tell for which one you want is what your changes sound like out loud. Put this clip here for exactly this long is a timeline sentence.

Cover this phrase with something. Delete this line. Make it faster. Those are transcript sentences, and they are the ones this shape handles.

The trade is deliberate. The machine owns the time arithmetic, so it rebuilds in seconds. Keep any of it yourself and it cannot.

There is a learning curve. It is short and specific. For about two days your hands reach for controls that are not there. Then you stop noticing they were.

  • Highlight a phrase and a clip lands over exactly those words.
  • Swap the clip in any spot, or search millions of free ones.
  • Delete a line you do not want and the cut rebuilds around it.
  • Change the pace and the machine re-cuts the whole thing.
  • Emphasise a word so it pops in the captions, or fix a mis-heard one.
  • Change how each clip enters: Cut, Whip, Punch, Glitch or Sweep.

This is the whole editor

Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.

CLIPS · 5I have thisexact conversationeverysingle week. Somebody sits down and says,oh yeah, I takecinnamonevery day.And honestly, doc, I have no idea if it works.So let me tell you what isin that capsule.a clip lands on these wordscut from the editTAKING CINNAMONEVERY DAY?is it doing anythingHeadlineMusicCaptionsTHIS VIDEOLength25.0sClips5Words removed18Export video

The gain is not speed on one video, it is the twelfth one existing

A single video is not much faster this way. Anybody claiming otherwise is comparing a rushed timeline edit to a careful one.

The change is that the second, sixth and twelfth cost the same as the first. On a timeline the twelfth is worse than the first. Attention depletes, and nobody has found a way around that.

Published benchmarks put the winner share at 5 to 8 percent, per Motion's analysis of 550,000+ Meta ads. That is roughly fourteen tested creatives per expected winner.

Six a month is not a creative problem. It is an under-sampling problem. Assembly is the one line in the cost stack with a fixed rate available.

In Cutroom a batch is 100 credits. An export is 20 credits per output minute. A thirty second ad is 110 credits, against the 2,500 Basic carries monthly.

Read that arithmetic before deciding any of this matters to you. A published range is a range rather than a promise. Your own win rate is the number worth logging.

Minutes of your attention, by position in the month

The scene most people recognise. The third bar is the one that decides whether week four has anything new to test.

  • First ad of the month45 careful minutes
  • Sixth ad20 and a shrug
  • Twelfth adnever got made

Four jobs stay on a timeline. The fifth decides your month.

Four jobs are timeline work and always will be. Naming them takes less time than discovering them.

Music-led edits where cuts land on a beat. There is no speech to index, and the placement is frame-level work.

Comedic timing, where the joke needs two extra beats held before the cut. Word level is not fine enough for that.

Multi-speaker interviews and montages, where several sources interleave. One continuous take is the assumption underneath everything here.

Anything that is not a vertical ad. The output is a 9:16 MP4 and the source ceiling is three minutes. What leaves the building is the finished file.

The fifth job is the one that decides most advertising months. The twelfth vertical ad. On a timeline it costs what the first did, so it usually never gets made. Here it costs a batch and an export.

That is the only claim worth making about the shape. It does not edit better than a timeline. It makes the ad that would not otherwise exist.

A batch is 100 credits and an export is 20 credits per output minute, so the twelfth thirty second ad costs 110. From a take you already uploaded it costs 10.

Run the arithmetic against your own month. If your count of tested creatives goes up, the shape did its job. If it does not, nothing else about the tool matters.

Which shape fits the job

Four rows name the jobs a timeline owns. The last row is the one that decides how many ads your month produces.

TimelineTranscript
Cuts landing on a musical beatThe right toolNo speech to index
Holding two extra beats for a jokeFrame level controlWord level, and no finer
Several speakers interleavedAny number of sourcesOne continuous take
Square and horizontal outputAny shape9:16 MP4 only
The twelfth ad this monthCosts what the first didA batch and an export

Questions people ask

What happens if the transcript mishears a word?
You fix it in the transcript and everything downstream follows, because the captions and the cut both read from the same text. Product names and unusual spellings are the usual offenders, and correcting them once per project is a normal part of the loop.
Can I remove a single word rather than a whole line?
The unit you act on is a line, or the pace lever, which re-cuts everything. That is a real constraint rather than an oversight: keeping the operations coarse is what lets the machine rebuild the sequence reliably every time.
Is this the same as text-based editing in a normal editor?
Related, and the difference is what happens after the text changes. In a timeline tool the transcript is a selection aid over a track you still own. Here the text is the only surface, and there is no track underneath to fall back to.
Is this the right shape for every edit?
It is built for vertical ads at volume. Music-led edits, comedic timing, multi-speaker work and output that is not vertical stay on a timeline. Say your next change out loud. If it names a sentence rather than a clip and a duration, this is the shape you want.

Name a sentence and the whole cut rebuilds in seconds. That is the gesture Cutroom is built on, and it is why the twelfth ad of the month costs what the first did.

Start with one take300 free credits · no card · cancel anytime