← All posts

Talking Head Videos: How to Make One Worth Watching

Why talking head videos lose viewers, the change interval that holds attention, and how to make a single camera setup watchable without a second angle.

Why a person speaking to camera loses viewers, how often the picture needs to change to hold them, and how to make a single camera setup watchable when a second angle is not an option.

A talking head video is one person speaking directly to camera, and it fails for a predictable reason: the picture stops changing. The consistent advice across production guides is to vary what the viewer sees at regular intervals, whether through a second angle, b-roll, a shot size change or on screen text.

The useful question is not "how do I make this more interesting?" It is "how long has the picture been identical?" Attention does not drift because the subject is boring. It drifts because nothing has changed for forty seconds, and that is a production problem with production fixes.

__wf_reserved_inherit

Why they lose people

A talking head is the least visually varied format in video. The frame is fixed, the subject barely moves, the background never changes, and after a while the eye has nothing left to do.

This is different from the content being weak. A genuinely interesting person saying genuinely interesting things still loses viewers to a static frame, which is why broadcast interviews have been cut between angles for decades rather than left on one camera.

Everything below is a way of giving the eye somewhere to go. The techniques matter less than the frequency.

Set the frame properly first

Most talking head problems start before anyone presses record.

Put the camera at eye level. Below eye level makes the subject look imposing, above makes them look diminished, and a laptop on a desk gives you the second one by default. Raise it.

Leave clear headroom and no more. A gap of empty wall above the head reads as carelessness. A crop through the top of the head reads as an accident.

Frame slightly off centre. Dead centre is static. Placing the subject a little to one side with space on the side they are looking toward gives the frame some direction.

Put depth behind them. The single biggest difference between a professional looking talking head and a domestic one is distance between subject and wall. Two metres of separation and a background thrown out of focus does more than any camera upgrade.

Simplify the background. A shelf with two objects reads as considered. A shelf with twenty reads as clutter, and the viewer will spend the first ten seconds reading it instead of listening.

Frame for the crop you will need. If clips will come out of this, keep the subject framed so a vertical crop still contains a whole head. That framing decision is covered in how to make shareable podcast clips for social media.

Light it so the face carries

One soft source in front, slightly off to one side. Large and soft beats bright and small. A window works if it is not behind the subject.

Never backlight. A window behind the head produces a silhouette and there is no fix in post.

Separate the subject from the background. A light on the background, or simply more distance, stops everything merging into one plane.

Do not move anything mid recording. If the light changes between takes, the cuts become visible and you have created a problem in the edit.

Change what the viewer sees

This is the part that matters most, and the specific method is less important than the interval.

A second camera angle is the strongest option. Cutting between a wide and a tighter shot of the same person is invisible to the viewer and gives the editor somewhere to cut every time speech is trimmed. Even a phone on a tripod as the second angle transforms the edit. The reason it works, and the editing sequence around it, is in how to edit an interview video.

B-roll covers both problems at once. It gives the eye somewhere new and hides the joins where filler was removed. Covered in how to clip multi speaker panel discussions without losing context.

A scale change on the same footage works with one camera. Punching in ten or fifteen percent across a cut reads as an edit rather than a glitch. It is a compromise and it is available to everybody.

On screen text counts as a change. A key phrase appearing as the speaker says it resets attention, and it also serves the large share of viewers watching muted. The treatment choices are covered in caption styles for B2B and B2C platforms.

Move yourself occasionally. Gesture, lean, turn slightly. A speaker who physically does not move for two minutes is harder to watch than one who does, independent of what they are saying.

__wf_reserved_inherit

Deliver it like a conversation

Speak to one person, not an audience. "You" rather than "you guys". The format is intimate and treating it as a broadcast makes it cold.

Vary the pace deliberately. A flat delivery at constant speed is soporific regardless of content. Slow down on the important sentence and speed up through the setup.

Use pauses as punctuation. A beat before the key point does more than emphasis does, and it also gives an editor a clean place to cut.

Look at the lens, not at yourself. Watching your own preview is the most common reason a speaker reads as distracted.

Plan the structure, not the wording. A scripted read sounds like a read. Bullet points keep you on track while leaving the delivery natural.

Keep it short, then cut it shorter

The consistent recommendation across guides is around two to three minutes for a standalone talking head, and shorter is usually better.

The reason is structural. A talking head has no subplot, no location change and no second character. The format cannot sustain the length that other video formats can, and nothing about production quality changes that.

The most reliable way to tighten one is to cut the filler in the middle rather than the ends. Removing hesitations and false starts through the body of the piece takes a three minute take to two minutes twenty without losing a single point, which is faster and better than trimming the introduction repeatedly.

Worked example: the same take, two edits

This is an invented example, not measured data.

A founder records a two minute explanation in one take on one camera.

The first edit trims the start and end and publishes it. It is one fixed frame for a hundred and ten seconds. The content is fine, and the retention graph falls steadily from the first ten seconds, because nothing in the picture ever changes.

The second edit uses the same take. Filler is cut throughout, which creates six joins. Four are covered with a fifteen percent punch in, two with screen recordings of the thing being described, and a key phrase appears as text at forty seconds. Total runtime, one minute forty.

Nothing was reshot. The picture now changes roughly every fifteen seconds and the piece holds. The production value did not improve, the change interval did.

Where Montage fits

Montage does not film anything, does not light anything and has no role in how you set up a talking head. Everything above happens before it is involved, and it will not add b-roll or cut between angles for you.

What it does is take a long recording and return the moments worth publishing. Upload up to 20GB at 4K and it gives you 8 to 10 scored clip candidates rather than an unranked pile, each trimmable by editing transcript text, with branded captions and vertical reframing applied, exporting as MP4 for social or XML, FCPXML and JSON for an editor.

Two things about that are relevant to a talking head specifically. The captions provide one of the change mechanisms described above without any extra work. And the XML or FCPXML export hands an editor a timeline with the cuts already made, so covering the joins happens on a sequence rather than from scratch, which is described in how XML export solves the producer and editor handoff problem.

Limitations and troubleshooting

The video looks flat and amateur despite a good camera. Usually the background is too close. Move the subject forward and let the wall fall out of focus.

The subject is silhouetted. There is a window behind them. Turn the setup around, because no grade recovers a blown background and an underexposed face.

Retention falls steadily from the start. The picture is not changing. Add a change every fifteen to twenty seconds by whatever means is available.

Every cut is visible. Single camera with no coverage. Use a scale change across each join, or accept a consistent visible rhythm rather than mixing hidden and obvious cuts.

The speaker looks stiff. They are reading. Switch to bullet points and record a second take without the script.

The clips cut from it will not crop to vertical. The frame was too wide. Reframe tighter at the shoot, since no tool adds picture that was never captured.

Frequently asked questions

What is a talking head video?

A video of one person speaking directly to camera, with no other subject or location. It is the most common business and creator video format and the least visually varied.

How long should a talking head video be?

Around two to three minutes for a standalone piece, and shorter where possible. The format has no subplot or location change to sustain a longer runtime.

How do you make a talking head video more engaging?

Change what the viewer sees at regular intervals, roughly every fifteen to twenty seconds. A second camera angle, b-roll, a scale change or on screen text all achieve it.

Do you need two cameras for a talking head video?

No, but a second angle is the single strongest improvement available and a phone on a tripod is enough. Without one, a fifteen percent punch in across each cut serves a similar purpose.

Why do talking head videos lose viewers?

Because the picture stops changing rather than because the content is weak. A fixed frame with a barely moving subject gives the eye nothing to do after the first thirty seconds.

Does Montage film or edit talking head videos?

No. Montage produces clips from recordings you supply. Filming, lighting and timeline editing all sit outside what it does.

Talking Head Videos: How to Make One Worth… | Montage Blog