How to Edit an Interview Video: Build the Story Before You Polish the Cut
A five-step interview editing workflow for building a truthful story spine, tightening answers, choosing useful B-roll, correcting captions and checking the final export.
You have a strong guest, a long interview and a deadline. A reliable way to make the interview engaging is not to start with jump cuts, captions or B-roll. First build a truth-preserving story spine from the guest’s complete answers; then tighten the dialogue, cover necessary cuts and finish the sound and picture.
That order matters. A polished interview can still feel empty when the editor has removed the hesitation, context or consequence that made the answer worth hearing.
The rule: cut for meaning before you cut for pace
An interview is not engaging because it moves quickly. It is engaging because the viewer can follow a question that matters, feel the tension inside the answer and reach a payoff that changes what they understand.
Your first edit should therefore answer four questions:
- What does this interview promise the viewer?
- Which answer carries the central tension?
- What context must stay for that answer to remain true?
- Where does the story resolve?
PBS’s editorial guidance gives the hard boundary: editing should never alter the meaning of a person’s remarks, and a cut should preserve enough context that the audience is not misled.[1] That standard is useful beyond journalism. A customer interview, founder story or expert conversation loses trust when a cleaner edit makes the speaker sound more certain, more dramatic or more absolute than they were.
A five-step interview editing workflow
1. Log the footage and write a one-sentence promise
Before moving clips, watch or read the full interview. Mark the questions, complete answers, strong examples, claims that need checking, emotional turns and any footage problems. Do not select only the sentences that sound impressive in isolation.
Then write one sentence:
This interview helps [viewer] understand [problem or change] through [guest’s experience or evidence].
For example:
This interview helps first-time product leaders understand why a growing roadmap can hide a broken setup experience through one founder’s failed launch and correction.
That sentence is your rejection tool. A funny tangent may be good, but if it does not deepen the promise, save it for another cut.
BBC documentary editors describe doing a paper edit from the interview words, voice-over and pieces to camera before the full picture edit.[2] You can make a lighter version with five columns:
| Source range | What the guest says | Story job | Must-keep context | Verification |
|---|---|---|---|---|
| 06:40–08:10 | Launch looked successful at first | Setup | Define “successful” | Check signup figure |
| 14:05–16:20 | Customers abandoned setup | Tension | Keep the observed behavior | Product screen cleared? |
| 28:30–30:15 | Team removed two steps | Decision | Preserve the qualifier | Confirm chronology |
| 42:00–43:10 | Activation improved later | Result | Avoid causal overclaim | Source for result |
The paper edit lets you test the argument before you spend time hiding cuts.
2. Build the story spine from complete answers
Move selected answers into story order. The interview does not have to preserve recording order, but every move should preserve meaning, chronology and the relationship between question and answer.
A reliable structure is:
- Open: the sharpest complete answer or moment of consequence.
- Context: who the guest is and what was happening.
- Tension: the mistaken belief, conflict or unresolved question.
- Evidence: the example that makes the tension credible.
- Decision: what the guest did or concluded.
- Resolution: what changed, what remains uncertain or what the viewer should carry forward.
Keep the question when the answer depends on it. If the guest says, “That was the point we stopped,” the viewer needs to know what that means and what stopped. Do not place an answer beside a different question or combine fragments into a sentence the guest never said; PBS specifically warns against both practices.[1]
Play this radio edit with the screen hidden. If the story is confusing without pictures, B-roll will only conceal the structural problem.
3. Tighten the dialogue without flattening the person
Now remove repetition, false starts and detours that do not change the meaning. Work answer by answer.
For each proposed cut, ask:
- Did I remove a qualifier such as sometimes, in this case or we think?
- Did the pause show uncertainty, emotion or a change of mind?
- Does the next sentence still follow logically?
- Would the guest recognize this as a fair version of the answer?
A clean cut does not need to sound machine-perfect. Leave breaths, laughter and short pauses when they carry personality or weight. If a statement needs heavy reconstruction to work, use another answer or paraphrase it in narration instead of manufacturing a quote.
Use a visible or audible cue when you jump across separate parts of an answer and seamless continuity would mislead. A different angle, a clearly relevant cutaway or an honest hard cut can show that time has moved.
4. Cover cuts with pictures that add information
Only after the dialogue cut works should you choose B-roll, alternate angles, graphics and reaction shots.
Adobe defines B-roll as supplemental footage and notes that it can establish a scene, smooth transitions and cover unwanted frames.[3] The key word is supplemental. Use it to show the product, place, person or evidence being discussed, not as decoration.
Use this hierarchy:
- Evidence: the screen, object, document or action the speaker names.
- Orientation: where the interview happens or who is involved.
- Reaction: a genuine response captured at the relevant moment.
- Continuity: an alternate angle or cutaway that honestly covers a dialogue edit.
Never use an unrelated reaction shot to imply an emotion that happened elsewhere. Do not let B-roll make two separate statements look like one seamless quotation.
For conversational flow, use split edits deliberately. In a J-cut, the next shot’s audio starts before its picture; in an L-cut, the previous audio continues after the picture changes.[4] These can move the viewer into a new thought or let a real reaction breathe. They should clarify the story, not announce the editor.
5. Finish sound, captions, graphics and export
Lock the story before detailed polish. Then:
- balance dialogue so speaker changes do not force volume adjustments;
- reduce distracting noise without making voices sound processed;
- add music only where it supports the tone and stays below speech;
- correct color and exposure between angles;
- add names, titles and graphics once, using verified spelling;
- generate captions from the final cut, then review every line;
- watch the exported file from beginning to end.
W3C guidance says captions should include the meaningful audio needed to understand prerecorded media, including speaker identification and relevant non-speech sounds.[5] Check names, technical terms, punctuation, speaker changes and whether the captions cover important visuals.
For YouTube delivery, the current recommended upload settings include H.264 video and a 48 kHz audio sample rate.[6] Treat destination settings as a final technical check, not a storytelling decision.
Worked example: the tighter cut is not always the better cut
This example is invented to demonstrate the method; it is not customer data.
A 55-minute customer interview contains this answer:
“We thought the launch worked because signups doubled. Three weeks later, support showed us that most new teams never invited a colleague. We removed two setup steps and changed the first project from a blank page to a guided example. That was when people started reaching the part of the product we had built the launch around.”
Two edits are possible:
Cut A: “Signups doubled. We removed two setup steps, and people started reaching the product.”
Cut B: Keeps the apparent success, the contradictory support evidence, the specific correction and the observed outcome.
Cut A is shorter but changes the story from we misread the launch and corrected the setup to we launched, changed two steps and succeeded. Cut B carries tension and causality without claiming more than the source says.
Build the spine around Cut B:
- open on “We thought the launch worked”;
- establish what the team measured;
- reveal the colleague-invite failure;
- show the guided example while the guest explains the correction;
- end on the observed behavior, not a generic success claim;
- flag the signup and activation statements for fact-checking.
That is an engaging edit because the viewer can follow a mistaken belief, evidence and decision. The engagement comes from the change in understanding, not the number of cuts.
Where Montage fits
Montage is useful when you have a spoken interview and want to direct one editable Moment from your footage. Current Moment creation is limited to under approximately three minutes per unit, so it does not produce a full-interview cut or a 30-minute edit.[9] Its current app-and-connector documentation lists Google Drive or public-URL import and AI-generated chapters for review.[8] Montage also supports MP4, XML, FCPXML and JSON outputs for publishing or an editor handoff.[7]
Chapters are a starting structure, not a finished story. A producer still chooses the promise, checks the source, protects the guest’s meaning, approves the picture and makes the final cut.
If your source is a full podcast episode, the same story-first principle applies; use this audio-and-video podcast editing workflow for the longer sequence.
Limitations and troubleshooting
The interview is accurate but dull. Look for a decision, reversal or unresolved question. If none exists, shorten the piece or change the format instead of adding more effects.
Every answer needs the full question. Add a concise setup in narration or on screen, but do not rewrite the question so it changes the answer’s meaning.
The cut has too many jump cuts. Restore natural pauses, use a second angle or add relevant B-roll. Do not cover a misleading composite quote with pretty footage.
The guest rambles. Select one complete thought, retain essential qualifiers and move secondary detail into narration or a separate clip.
The first cut is over the target length. Remove duplicated evidence and sections that perform the same story job before shaving words from the central answer.
Frequently asked questions
How do I make an interview video more engaging?
Give it one clear promise, build a story spine from complete answers, preserve tension and consequence, then use pictures and sound to clarify that story. More cuts do not automatically create more interest.
Should I remove every pause and filler word?
No. Remove what blocks comprehension or pace, but keep pauses, breaths and hesitations when they carry emotion, uncertainty or personality.
Can I change the order of interview answers?
Yes, when the new order remains faithful to meaning and chronology. Do not pair an answer with a different question, hide a material time jump or combine fragments into a statement the guest never made.[1]
What B-roll should I use in an interview?
Prefer footage that proves or explains what the speaker says: the object, place, person, action, screen or document in question. Use alternate angles and cutaways for continuity only when they do not imply a false reaction or timeline.[3]
Should I edit an interview from the transcript or the timeline?
Use both. A transcript or paper edit helps you see repetition, themes and structure; the picture and sound reveal delivery, body language, timing and visual continuity that text cannot show.[2]
Build the honest story first
The repeatable rule is simple: decide what the interview means before deciding how polished it should look. Keep the context that makes the answer true, make each section perform one story job and let the finish support the speaker rather than rewrite them.
If you have a long interview and know the short Moment you want to produce, use Montage to direct that editable Moment from your footage. You keep the taste and final judgment.
Sources
- https://www.pbs.org/standards/blogs/standards-memos/handling-quotations
- http://downloads.bbc.co.uk/academy/academyfiles/080617%20How%20to%20edit%20a%20documentary%20%28Transcript%29.pdf
- https://www.adobe.com/creativecloud/video/discover/b-roll.html
- https://www.adobe.com/creativecloud/video/discover/j-cut-and-l-cut.html
- https://www.w3.org/WAI/WCAG22/Understanding/captions-prerecorded.html
- https://support.google.com/youtube/answer/1722171
- https://montage.app
- https://montage.app/mcp
- https://montage.app/llms.txt