How to Edit a Podcast: An Audio and Video Workflow
A practical seven-step workflow for editing a video podcast: sync the tracks, lock the story, repair cuts, mix the audio, polish the picture, review the full episode, and export each delivery format.
Editing a video podcast works best as one locked sequence: organize and sync the source tracks, make the story cut, repair sentence-level joins, mix the audio, polish the picture, watch the full episode, then export separate files for the podcast feed and video platform. Do those jobs out of order and you create rework. Polish a camera angle before the story is locked, and the next content cut can wipe out that work.
This guide assumes you have permission to use the recording and an editor that can cut audio and video, add captions, and export delivery files.
Before you edit: define the deliverables
Write down what must leave the edit before you open the timeline. For a typical video podcast, that means:
- one full-resolution project master;
- one audio file for the RSS feed;
- one video file for the video platform;
- one checked caption file or burned-in caption version, if required;
- show notes, title, thumbnail, and short clips as separate downstream jobs.
Keep the full episode and its promotional clips as different outputs. Lock the episode master first so a late change does not leave every clip out of date. Preserve the original files and keep isolated tracks when available.
1. Import, label, and synchronize every track
Import the source media and label tracks by role: host audio, guest audio, host camera, guest camera, screen share, music, and graphics. Do not leave five files named after camera serial numbers on the timeline.
Synchronize the tracks before making content decisions. A hand clap, slate, or other sharp waveform peak gives you a visual reference. Automatic waveform sync is a useful starting point, but check it at the beginning, middle, and end of the episode. If mouths and words drift apart later in the recording, a single start-point alignment did not solve the whole file.
Once sync is stable, link each speaker's audio and video. Duplicate the synchronized sequence before making the story cut; the untouched version is your recovery point.
2. Make the story cut before polishing anything
Listen through the episode at normal speed. Mark three kinds of material:
- Keep: ideas, explanations, stories, and exchanges that move the episode forward.
- Maybe: tangents or repeated points that may be useful if the episode runs short.
- Remove: false starts, technical interruptions, dead setup time, and repetitions that add no meaning.
Cut for comprehension first, duration second. Removing every breath and pause can make a real conversation feel mechanically assembled. A short pause before an important answer may be part of the answer.
Log any edit that changes meaning: source timecode, what you removed, and why.
At the end of this pass, you should have one continuous episode that makes sense even if the picture is still rough and the levels are uneven.
3. Repair sentence cuts and transitions
Now revisit every edit point. Listen from a few seconds before the cut to a few seconds after it. Riverside similarly tells editors to listen back after removing a word or section and to smooth any abrupt jump in the timeline.[3]
Check each join for:
- a word clipped at the beginning or end;
- a breath that starts but does not finish;
- a sudden change in room tone;
- a jump in speaking rhythm;
- a reaction shot that no longer matches the dialogue;
- a sentence whose meaning changed when the surrounding context disappeared.
Use room tone to bridge tiny gaps when necessary. If a cut still sounds unnatural, restore a little of the pause or move the edit to a cleaner phrase boundary. The goal is not the smallest possible gap. It is a sentence the listener never has to decode twice.
4. Repair and mix the audio
Work from the dialogue outward. Mute unused microphones where they only add room noise or bleed. Remove obvious clicks, hum, and distracting background noise conservatively. Heavy processing can replace one problem with metallic speech or clipped consonants.
Then balance the speakers. Use clip gain or automation for large level differences before asking compression to do everything. Add music and other sound only after the dialogue remains clear on its own. Duck the music under speech and check the result on headphones and ordinary laptop or phone speakers.
For Apple Podcasts, Apple recommends preparing podcast audio around -16 dB LKFS, with a tolerance of ±1 dB, and keeping true peak at or below -1 dB FS.[1] Treat that as an Apple delivery target, not a universal law. Confirm the specification for your hosting platform and audience before export.
Do a final noise-floor and level check after the content edit is locked.
5. Polish the video after the audio story is locked
Choose camera angles to clarify who is speaking and what the audience needs to see. Do not switch angles merely because a fixed number of seconds has passed. A held reaction can be useful; a rapid sequence of unnecessary cuts can distract from the conversation.
Check:
- speaker framing and eye line;
- crop consistency across cameras;
- screen shares at a readable size;
- exposure and white-balance differences;
- jump cuts that need a second angle, punch-in, or clean hold;
- logos, lower thirds, and episode titles;
- captions against names, specialist terms, and numbers.
Review captions as editorial copy. Correct spelling, punctuation, speaker changes, and timing. Prepare a separate caption file when the destination supports it.
6. Watch the full episode in one sitting
A timeline review is not a viewer review. Export a review file and watch it from the first frame to the last without editing as you go. Log problems, then return to the project once the playback ends.
Run four checks:
- Story: Does the opening establish the subject quickly? Does any removed context make a later reference confusing?
- Sound: Are both speakers intelligible at the same listening volume? Are music and stings lower than dialogue?
- Picture: Is sync stable? Are camera changes, graphics, and captions correct?
- Delivery: Are the episode title, version, duration, and filename unambiguous?
After fixes, watch the changed sections with generous handles on both sides. For a high-stakes release, do another full playback.
7. Export separate audio and video files
Do not treat one compressed video file as the master for every destination. Export a high-quality archive or project master, then make destination-specific deliveries from that master.
Apple Podcasts accepts MP3 or AAC for RSS-feed audio. Its published ranges for 44.1 or 48 kHz audio are 64–128 kbps for mono and 128–256 kbps for stereo; Apple also notes that AAC gives better quality than MP3 at the same bit rate.[1] Your podcast host may impose its own limits, so check those before upload.
For YouTube, the current recommended settings include an MP4 container, H.264 video, 48 kHz audio, and the same frame rate used for the recording.[2] Those are YouTube recommendations, not defaults for every video destination.
Name exports so another person can identify the approved version without opening them. For example:
show-042-master-v03.movshow-042-rss-v03.mp3show-042-youtube-v03.mp4show-042-captions-v03.srt
Open every final file after export. Check its duration, first and last frames, audio channels, caption attachment, and filename. A completed export is not yet a quality check.
Worked example: a 52-minute two-person interview
Suppose the recording contains two isolated microphones, two cameras, and a mixed reference track.
First, the producer labels and synchronizes all five files, then checks sync near minute 5, minute 28, and minute 49. The story pass removes six minutes of setup, one duplicated answer, and a connection interruption. The episode now runs 44 minutes.
The transition pass restores half a second before one guest answer because the first cut sounded rushed. The mixer balances the two microphones, removes a short electrical buzz, and lowers the intro music under the host's opening. The video pass uses the guest camera for the key explanation and a two-shot for the exchange that follows. Captions are corrected for the guest's surname and two product names.
The producer watches the 44-minute review export in one sitting, logs three problems, fixes them, and exports a new master. From that approved master, they create the RSS audio, YouTube video, and caption file. Only then do they start making promotional clips.
This example shows the order of operations. It is not a claim about typical editing time, quality, or output volume.
Common problems and fixes
The voices drift out of sync
Check whether the source tracks were recorded at different frame rates or sample rates. Re-sync near the point where the drift becomes visible instead of stretching the whole sequence blindly. If the source itself is unstable, document the repair and check the end again.
A dialogue cut sounds artificial
Restore a little room tone or breathing space. Move the edit to a complete phrase. Listen with the picture hidden; visual movement can distract you from an audible jump.
Noise reduction makes speech metallic
Back off the processing and accept a small amount of steady room noise. Consistent, intelligible dialogue is usually less distracting than aggressively processed speech that changes from word to word.
The video export looks soft or jerky
Check the source resolution, sequence settings, and output frame rate. For YouTube, keep the upload at the same frame rate as the recording rather than converting it without a clear need.[2]
The captions disagree with the speaker
Correct the text against the audio, then review timing around every correction. Names, numbers, and technical terms deserve a manual pass even when the first transcript was generated automatically.
Where Montage fits, and where it does not
This workflow stands on its own. You can follow it in the editor you already use.
If you have an authorized video-podcast recording and want help directing editable video from your footage, Montage may fit that part of the workflow. You keep the taste and final judgment. Montage does not replace the audio master, the delivery-spec check, or the final playback.
If the part that slows you down is moving a first cut into Final Cut Pro, use this editor-handoff checklist.
Already have an authorized video-podcast recording? See whether Montage fits the way you want to direct the first cut.
Frequently asked questions
Should I edit podcast audio or video first?
Lock the spoken story first while audio and video remain synchronized. Then repair and mix the audio before doing detailed visual polish. That sequence prevents a late content cut from undoing finished picture work.
What loudness should a podcast be?
Apple Podcasts recommends about -16 dB LKFS, ±1 dB, with true peak no higher than -1 dB FS.[1] Confirm the requirement for your own host and destination rather than assuming every platform uses the same target.
Should a video podcast use the mixed call recording?
Use isolated speaker tracks when they are available and intact. Keep the mixed call recording as a reference or backup. Isolated tracks give you more control over noise, bleed, and speaker balance.
Can Montage edit an audio-only podcast?
No. This article uses Montage only as an optional fit for authorized video footage. Use an audio editor or digital audio workstation for an audio-only episode.
When should I make podcast clips?
After the full episode master is approved. That keeps clip boundaries, captions, and quotes tied to the final version rather than an earlier cut.
Sources
[1] Apple Podcasts — Audio requirements
[2] YouTube — Recommended upload encoding settings
[3] Riverside — How to Edit a Podcast: Audio & Video Guide (2026)