← All posts

Full Verbatim vs Clean Verbatim: Which Transcript Do You Need

Full verbatim, clean verbatim and edited transcription compared with side by side examples, when each is required, and what it costs to choose wrong.

The three transcription styles compared with the same passage shown in each, what every one of them keeps and discards, which fields require which, and why the choice cannot be deferred.

Full verbatim records every utterance including filler, repetition, false starts and non verbal sounds. Clean verbatim removes those while keeping the wording and meaning untouched. Edited transcription additionally smooths grammar and sentence structure, which makes it readable and means it is no longer a record of what was said.

The useful question is not "which style is most accurate?" All three are accurate at what they set out to capture. It is "does how this was said matter, or only what was said?" Answer that and the style chooses itself.

__wf_reserved_inherit

The same passage, three ways

This is the fastest way to understand the difference.

What the speaker actually said:

"So, I mean, the, the thing is we, we tried it for about, I want to say, six months? And, um, honestly it just, it didn't, it didn't move the number at all. Like at all. [laughs]"

Full verbatim. Everything, including the stutters, the filler, the upward inflection on "six months" and the laugh.

"So, I mean, the, the thing is we, we tried it for about, I want to say, six months? And, um, honestly it just, it didn't, it didn't move the number at all. Like at all. [laughs]"

Clean verbatim, also called intelligent verbatim. Filler, stutters and false starts removed. No word reordered, no grammar corrected, meaning untouched.

"The thing is we tried it for about six months, and honestly it didn't move the number at all. Like at all. [laughs]"

Edited transcript. Grammar tidied, repetition removed, sentence smoothed for a reader.

"We tried it for about six months and it did not move the number at all."

What each version lost

Reading the three in sequence shows exactly what the trade is.

Clean verbatim lost the hesitation before "six months", which suggested the speaker was estimating rather than recalling, and the stutter on "it didn't", which in context read as frustration. For most purposes neither matters. For a researcher studying how people talk about failed initiatives, both are data.

The edited version lost considerably more. "Like at all" was the emphasis, and it is gone. The laugh is gone, which in the original was self deprecating and changed the tone of the whole statement. What remains is accurate in substance and no longer representative of the moment.

That last point is the one worth holding onto. An edited transcript is a legitimate document and an illegitimate quote. If you publish the edited version inside quotation marks, you have published something the person did not say.

Which style each field requires

Field Style Why
Legal, court, deposition Full verbatim Manner of speech is evidence. Hesitation, correction and tone can all be material
Academic qualitative research Full verbatim Pauses and filler are analysable data, and reviewers expect them
Market and UX research Clean verbatim Themes and the participant's own phrasing matter, hesitation usually does not
Journalism Clean verbatim Quotes must be the speaker's words, tidied only of filler
Medical documentation Full verbatim Regulatory and liability requirements
Podcast and video production Clean verbatim The transcript is a tool for locating moments
Show notes and website copy Edited Written for a reader, not presented as a quote
Subtitles and captions Clean verbatim Filler wastes reading time on screen

Two observations from that table.

Clean verbatim is the correct default. Four of the eight rows, and the overwhelming majority of transcription work, needs it. Full verbatim is required by a minority of fields and produces a document that then has to be cleaned before any ordinary use.

Only one row calls for edited, and in that row the output is not being presented as a transcript at all.

What counts as filler

The boundary between full and clean verbatim is a list of things one keeps and the other does not. Being specific about that list is what makes a transcript consistent.

Removed in clean verbatim:

  • Filler words: um, uh, er, ah, mm
  • Discourse fillers used as padding: like, you know, I mean, sort of, basically, actually
  • Stutters and repeated words: "the, the thing"
  • False starts: "we tried, we started by trying"
  • Stammering and self corrections mid word

Kept in clean verbatim:

  • Every substantive word
  • The speaker's grammar, including where it is non standard
  • Word order exactly as spoken
  • Meaningful non verbal sounds, usually bracketed, such as [laughs]
  • Regional and colloquial phrasing

The hard cases, where transcribers disagree and a project convention is needed:

  • "Like" and "you know" used meaningfully rather than as filler. "It was like a factory in there" is substantive. "It was, like, difficult" is filler.
  • Repetition used for emphasis. "It was bad, bad" may be deliberate.
  • Trailing off, where "I just thought..." carries meaning the completed sentence would not.

Decide these once at the start of a project and write the decision down. Consistency between transcripts matters more than which side you land on.

__wf_reserved_inherit

Cost, time and the practical consequence

The styles are not equally expensive to produce.

Full verbatim is slower in every method. Typing it means capturing sounds that are harder to hear than words. Correcting an automatic draft into full verbatim means adding back filler that the model frequently drops by default, which is tedious work most people underestimate. Human transcription services routinely charge more for it.

Clean verbatim is what automatic transcription naturally produces, roughly. Most models drop some filler on their own, which is convenient when you want clean verbatim and a problem when you need full.

That last point has a consequence people discover late. If you need full verbatim, automatic transcription is not a shortcut to it. You will be reinstating filler by ear, and at that point typing from scratch may be faster.

Converting between styles is one directional in practice. Full verbatim to clean verbatim is a deletion pass, which is mechanical and fast. Clean verbatim to full verbatim is not possible from the document at all, since the removed material is gone. You have to go back to the audio.

This is why the decision belongs before the work rather than after it.

How it affects transcripts used for video

If the transcript exists to locate moments in a recording rather than to document it, clean verbatim is correct and the reasoning is slightly different from the other fields.

A transcript used for cutting is read, not studied. Filler makes it slower to scan, and the point of reading a transcript instead of watching the footage is speed. Reading an hour of clean transcript takes about twelve minutes. Full verbatim of the same hour takes noticeably longer and surfaces nothing extra for this purpose.

Where it matters is the cut itself. A clip trimmed from a clean verbatim transcript will not include the "um" that preceded the sentence unless the tool cuts from the original audio rather than the cleaned text. That is a detail worth checking in any transcript driven editor, and the approach is described in text based video editing.

For chaptering and navigation, clean verbatim is also the right input, as covered in how to create video chapters from a transcript.

Captions are a separate decision

Caption text is effectively clean verbatim and more aggressive, because reading speed is the constraint.

On screen, a caption containing "um, I mean, the, the thing is" wastes the viewer's reading time on words carrying no information, and the text has to keep pace with speech. Captioning conventions therefore remove filler by default and frequently condense further.

That is a different standard again from a reading transcript, and the formats and when each applies are set out in captions vs subtitles vs SDH.

Worked example: one interview, two deliverables

This is an invented example, not measured data.

A researcher interviews twelve people for a study and the same recordings are also going to produce marketing content.

The study needs full verbatim. Hesitation before answering a question about cost is analysable, the reviewer expects it, and the convention has to be identical across all twelve transcripts.

The marketing team needs clean verbatim. They are looking for quotable statements and clip-worthy moments, and the filler slows the read without adding anything.

The efficient route is to produce full verbatim once, since it is the only direction that converts, then run a deletion pass to produce the clean version for the marketing team. Producing clean verbatim first would have meant transcribing all twelve recordings again for the study.

The order was the whole decision, and it was made before any transcription started.

Where Montage fits

Montage is not a transcription service and does not produce full verbatim transcripts. For research, legal or certified work requiring full verbatim with a documented convention, use a dedicated transcription provider.

It produces a clean verbatim style transcript as the first stage of turning a long recording into publishable video. Upload a recording, up to 20GB at 4K, and it returns a timestamped transcript alongside 8 to 10 scored clip candidates, each trimmable by editing the transcript text, with branded captions and vertical reframing applied, exporting as MP4 for social or XML, FCPXML and JSON for an editor.

That style suits the job it is doing, which is reading quickly to find moments, and it is the wrong deliverable for anything where filler is data. The pipeline behind the clip selection is explained in how AI video clipping works, and you can try it on a single recording with the podcast clip finder.

Limitations and troubleshooting

The automatic transcript dropped filler and I needed full verbatim. Most models clean by default. Reinstating filler by ear is slow enough that typing from the audio may be faster, and this is the main reason to decide the style before choosing the method.

Two transcripts in the same project treat "you know" differently. No project convention was written down. Decide the hard cases once, record the decision, and apply it to every transcript including the ones already done.

The quote reads awkwardly in print. That is clean verbatim doing its job. Either publish it as spoken, or mark it as edited for clarity, but do not silently smooth it inside quotation marks.

A researcher rejected the transcripts. Almost always clean verbatim supplied where full verbatim was required. The material cannot be recovered from the document, so the recordings have to be transcribed again.

Clean verbatim still reads as cluttered. Some speakers use substantive phrasing that looks like filler. Removing it crosses into edited territory, which changes what the document is.

Non verbal sounds are inconsistent. Decide which ones carry meaning for your project, usually laughter and long pauses, and ignore the rest rather than annotating every breath.

Frequently asked questions

What is the difference between verbatim and clean verbatim?

Full verbatim captures every utterance including filler, stutters, false starts and non verbal sounds. Clean verbatim removes those while keeping every substantive word, the original word order and the speaker's grammar intact.

What is intelligent verbatim transcription?

Another name for clean verbatim. The terms are used interchangeably and mean a transcript with filler and false starts removed but the wording otherwise unchanged.

When should you use full verbatim?

When how something was said is part of the evidence. Legal and court work, medical documentation, academic qualitative research, and any analysis where hesitation or repetition is itself data.

Is an edited transcript the same as clean verbatim?

No. Clean verbatim removes filler but does not change wording or grammar. An edited transcript smooths sentences for readability, which makes it unsuitable for presenting as a direct quote.

Can you convert clean verbatim back to full verbatim?

Not from the document, because the removed material no longer exists in it. You would have to transcribe again from the audio, which is why the style decision has to be made first.

Does automatic transcription produce verbatim or clean verbatim?

Closer to clean verbatim, since most models drop some filler by default. That is convenient when clean is what you want and a significant obstacle when you need full verbatim.

Which style should captions use?

Clean verbatim, and usually more condensed still, because reading speed limits how much text can be on screen while keeping pace with the speech.

Which style does Montage produce?

A clean verbatim style transcript, intended for reading quickly to locate moments in a recording. For full verbatim deliverables, use a dedicated transcription service.

Full Verbatim vs Clean Verbatim: Which Transcript… | Montage Blog