Feature
The ums go, and the transcript shows you every word that went
Cutroom's transcriber flags the hesitation sounds in your take, and the fast pace drops them from the cut. You audit the removals by reading rather than by scrubbing, because every dropped word stays visible in the transcript. Here is exactly what gets removed, what deliberately stays, and the one thing this will not do to your sentences.
- The pace that drops the flagged fillers
- fast
- To fix a word the machine heard wrong
- 1 click
- Words it rewrites or replaces
- 0
Flagged in transcription, dropped at the fast pace, visible the whole time
When your take is transcribed, every word gets a timing and the hesitation sounds get a flag. The ums, the uhs, the sounds that hold a gap open while a thought loads.
Set the pace to fast and the flagged fillers are removed from the cut, each one cut at exactly the span of the word itself, never a syllable more. The long pauses around them are tightened in the same pass.
At chill and normal, the fillers stay in. That is a decision rather than a gap, and the next section is the reasoning.
Nothing disappears silently. The transcript shows the removed words struck through, so you read what went the way you would read tracked changes. The audit takes seconds instead of a replay.
One um, from recording to gone
The flag is set during transcription. The pace decides what happens to it.
You say um
While the next thought loads
Transcriber flags it
During the 100 credit batch
Pace set to fast
Chill and normal keep it
Dropped at its exact span
Never a syllable more
Shown struck through
Audited by reading
Why chill keeps the ums on purpose
A hesitation is not always a flaw. In a relaxed, kitchen-table register, the ums are part of how trust sounds, and a slow talker with every hesitation surgically removed sounds dubbed.
So the pace setting is really a register setting. Fast is the direct-response register, where hesitation reads as drag, and the fillers go. Chill is the talking-to-a-friend register, where they carry warmth, and they stay.
Normal sits between the two. Pauses tightened, fillers kept.
That is also the honest answer to a question every filler remover dodges: which ums should go? There is no per-um checkbox here. You choose a register and the register decides consistently, which an ad needs more than it needs word-by-word control.
This is the whole editor
Highlight a phrase and a clip lands on those exact words. No timeline, no keyframes, no layers.
It removes fillers. It does not rewrite sentences
Here is the limit, stated plainly. If the sentence was clumsy, it is still clumsy without the um. No tool that works on your recorded voice can write you a better sentence, because the better sentence was never recorded.
What you can do is delete. Remove the line that went nowhere and the cut rebuilds around the gap, footage and captions included. One action, on the transcript.
What you cannot do is retype. Changing a spoken word into a different word would leave the audio disagreeing with the captions, so the editor does not offer it. The transcript edit that exists is a correction: fixing a word the machine misheard, which updates the caption and the footage search and never touches your audio.
If the take needs different sentences, record the take again. It is a phone and three minutes, and the second read is nearly always better anyway.
The by-hand version of this job, for contrast
Doing this on a timeline means finding each um on a waveform, razoring both sides, closing the gap, and playing the join to check it. A few minutes each, more when the um leans against the next word.
A three minute take from a normal speaker can carry a dozen of them. That is half an hour of the least creative work in editing, per take, per variant.
The transcript route is one setting. And because the cut is recomputed from your original take every time, switching between fast and normal to hear both registers costs nothing before export.
A dozen ums, two ways
The by-hand number repeats on every variant of the same take. The setting does not.
- One um, razored by hand3 min
- Twelve ums, by hand36 min
- Twelve ums, one pace settingunder 1 min
Cost, and who should use something else
The flagging and the removal happen inside the 100 credit batch, along with the cut, the captions and the footage search. Export is 20 credits per minute of finished video, so a thirty second ad is about 110 credits.
Lite is $19.99 a month for 1,250 credits, nothing on the picture, cancel in one click.
The scope is one spoken take of up to three minutes becoming one vertical ad. If you want the ums stripped from a lecture, a podcast or any file you need handed back afterwards, that is a different job. A long-form transcript editor is built for it. This one stops at three minutes and always returns an ad.
Questions people ask
- Which sounds count as filler words?
- The hesitation sounds the transcriber flags: the ums and the uhs. Spoken words that people lean on, a mid-sentence like or you know, are real words to the transcriber. If a line is full of them, the honest fix is deleting the line.
- Do I have to remove them?
- No. Fillers only drop at the fast pace. Chill and normal keep them, because in a relaxed register the hesitations are part of how a real person sounds.
- Can I remove one um but keep another?
- No. The pace decides consistently for the whole take. Per-um control sounds attractive and produces edits nobody can reproduce next week.
- The transcript got a word wrong. Now what?
- Click it and fix it. The correction updates the caption and what the footage search understands the line to mean. It never changes your audio, which stays exactly as recorded.
- Will it tighten my rambling sentence?
- It will not rewrite it. It removes fillers, shortens pauses, and lets you delete whole lines. A sentence that needs different words needs a new take, which is a phone and thirty seconds.
- What does it cost?
- Nothing on top. The flagging and removal ride inside the 100 credit batch, export is 20 credits per minute of finished video, and Lite is $19.99 a month for 1,250 credits, cancel in one click.
Say it with the ums in. Pick the register afterwards, read what went, and ship the take that sounds like you on a good day.