← All posts

How to Look and Sound Natural on Camera

The posture, eye line, breathing and delivery habits that make someone read as natural on camera, and the three mistakes that make people look stiff.

The specific physical and delivery habits that make someone read as natural on camera, why the eye line matters more than anything else, and the three things that reliably make people look stiff.

Looking natural on camera comes down to four things: look at the lens rather than at your own image, keep your body slightly in motion rather than held still, breathe properly before and during the take, and speak to one imagined person rather than to an audience. None of them require confidence, which is why they work before confidence arrives.

The useful question is not "how do I look more professional?" It is "what am I doing that I would not do in a conversation?" Natural is not a skill to acquire. It is the absence of the specific things people start doing the moment a camera is pointed at them.

__wf_reserved_inherit

The eye line decides most of it

This is the single highest leverage change available and it costs nothing.

Look at the lens, not at your own image on screen. A viewer experiences lens contact as eye contact. Looking at your preview means looking slightly below and to the side, which reads to a viewer as distraction or evasion even though neither is happening.

Most setups actively work against this, because the preview is more interesting than a small glass circle. Three fixes:

Cover or minimise your own preview. A sticky note over your own tile removes the temptation entirely.

Put the lens at eye level. A laptop on a desk puts the camera below your eye line, which means you are looking down at the viewer and they are looking up your face. Raise the device until the lens is level with your eyes.

Place anything you need to read immediately beside the lens, at the same height. The further your notes sit from the camera, the more obvious the eye movement.

One nuance that separates natural from robotic. You do not need unbroken lens contact. People in conversation look away while thinking and return when speaking. Letting your eyes drift upward or sideways mid thought and returning to the lens for the point reads as genuine recall. Staring fixedly for ninety seconds reads as a hostage video.

Three things that make people look stiff

Holding still. The most common and the most fixable. Nervous people lock their body, and stillness reads as tension regardless of what the face is doing. Gesture with your hands, shift your weight, lean slightly forward into a point. If your hands are below frame and never move, bring them up.

Shallow breathing. Anxiety shortens the breath, a short breath tightens the voice and raises its pitch, and a tight high voice sounds nervous, which makes you more nervous. Two or three slow breaths before the take, in through the nose and out through the mouth, interrupts the loop before it starts.

Performing rather than talking. A voice that is louder, faster and flatter than your normal register. It happens automatically when the mental frame is "I am making a video" rather than "I am explaining this to someone." The fix is the frame, not the voice.

Posture, and what actually reads

Sit or stand straight with shoulders back and chin level. Not rigid. The difference between upright and stiff is tension, not alignment.

Lean slightly forward rather than back. Forward reads as engaged and interested. Back reads as detached or sceptical, which is rarely what anyone intends.

Keep your hands visible and in use. Hands below the frame do nothing for you. Hands that move while you talk are the clearest signal that you are talking rather than reciting.

Do not fold your arms. An obvious one that people do anyway when nervous.

Settle before you start. Press record, breathe, then begin. The first few seconds of most recordings show someone still arranging themselves, and those seconds are usually cut, which makes it a free problem to have.

The framing side of all this, including where to put the camera and how much headroom to leave, is in talking head videos.

__wf_reserved_inherit

Speak to one person

The most effective mental change available, and it is purely a framing exercise.

Pick a specific individual. Not "my audience". A particular person you find easy to talk to, ideally someone who would actually want to know this. Imagine them behind the lens.

Use "you" rather than "you guys" or "everyone". Short form video is watched by one person holding a phone, and addressing a crowd breaks that.

Say it the way you would say it to them. If you would not use a word in that conversation, do not use it on camera. Corporate vocabulary is one of the clearest tells that someone has switched into performance mode.

This single change does more for naturalness than any amount of practice at "being natural", because it replaces the task rather than improving your execution of it.

Delivery habits worth building

Vary your pace deliberately. Slow down on the point, speed up through the setup. A constant rate reads as recitation no matter how good the content is.

Use pauses as punctuation. A beat before an important sentence does more than emphasis does. It also gives an editor a clean cut point, which matters more than people realise.

Finish sentences. Trailing off is a nervous habit and it makes passages impossible to use as standalone clips.

Let the first take be bad. Say the opening twice and keep the second. Almost everyone is stiffer in the first thirty seconds than they are afterwards, so plan to discard them.

Restate the question inside the answer when being interviewed. It makes you sound composed and it makes the answer usable on its own, as described in 60 podcast interview questions that produce clippable answers.

Why this matters for clips specifically

Short form is unforgiving about all of the above in a way long form is not.

A forty second clip has nowhere to hide. There is no warm up period, no later section where you relax into it, and no surrounding context to carry a flat delivery. The first three seconds are doing the work of an entire introduction, which is the argument in how to write a hook for short form video.

Two practical consequences.

Record the important points early. Delivery degrades over a session, noticeably after about ninety minutes. Putting the material you most want as clips at the front rather than working through a list in order is a meaningful difference, and it fits the approach in how to batch a month of clips in one afternoon

Leave a beat either side of each point. Clean silence around a moment is what makes it extractable. Overlapping or rushing between points is the most common reason a good passage cannot be cut cleanly.

Worked example: the same person, two takes

This is an invented example, not measured data.

A founder records an explanation of their pricing twice on the same afternoon.

The first take is done on a laptop on a desk, with their own preview visible. They read from a script taped beside the screen, sit still, and get through it in one pass. Watching it back, they look like they are reading, because they are, and their eye line drifts left on every sentence.

The second take is done with the laptop raised on a box so the lens is at eye level, the preview covered with a sticky note, and three bullet points instead of a script. They record it three times and keep the third. Their hands move, the delivery varies, and they look at the lens.

Nothing changed about the person, the room, the camera or the lighting. Four small adjustments and a decision to use multiple takes produced the difference.

Where Montage fits

Montage does not record, does not coach delivery and has no part in how you perform on camera. All of this happens before it is involved, and no tool compensates for a delivery that did not land.

Where it is relevant is the multiple takes point. Upload a recording, up to 20GB at 4K, and it returns a timestamped transcript alongside 8 to 10 scored clip candidates, each trimmable by editing the transcript text, with branded captions and vertical reframing applied, exporting as MP4 for social or XML, FCPXML and JSON for an editor.

Because the trimming happens on text, keeping the third attempt at a sentence and deleting the first two is a matter of deleting words rather than hunting for edit points in a waveform, which is described in text based video editing. Recording loosely and selecting afterwards is a reasonable way to work rather than a sign of weak delivery.

To try it on one recording, use the podcast clip finder.

Limitations and troubleshooting

I look like I am reading because I am reading. Move the text immediately beside the lens, narrow it to a few words per line, and switch from a full script to bullets.

My eye line is slightly off in every take. The lens is not at eye level, or you are watching your own preview. Raise the camera and cover the preview.

I sound higher pitched than normal. Shallow breathing. Breathe before the take and pause more often during it.

I look frozen from the chest down. Your hands are out of frame. Raise them into shot and let them move.

The first minute is always worse. Normal and universally true. Record a throwaway minute first and start the real content after it.

I am natural in conversation and stiff on camera. The mental frame has switched to performance. Pick one person and explain it to them rather than presenting it.

My delivery is fine and the clips still feel flat. Check whether you are leaving silence around your points. Rushing between them removes the space an editor needs and flattens the pacing.

Frequently asked questions

How do you look natural on camera?

Look at the lens rather than at your own preview, keep your body slightly in motion, breathe before and during the take, and speak to one imagined person rather than to an audience.

Should you look at the camera or at yourself on screen?

At the camera. A viewer reads lens contact as eye contact, while looking at your own preview reads as looking slightly away, which comes across as distracted even when you are concentrating.

Do you have to look at the lens the whole time?

No, and doing so looks unnatural. Letting your eyes drift while thinking and returning to the lens on the point mirrors how people behave in conversation.

Why do I look stiff on camera?

Usually three things together: holding your body still, breathing shallowly, and shifting into a performing voice. All three are habits that appear when a camera is pointed at you and are absent in ordinary conversation.

How do you sound more natural when speaking to camera?

Vary your pace, pause as punctuation, use the words you would use with a colleague, and work from bullet points rather than a word for word script.

Where should the camera be positioned?

At eye level. A laptop on a desk places the lens below your eyes, which means you look down at the viewer and they see up your face.

Does Montage improve how I come across on camera?

No. Montage processes recordings into clips. It is relevant only in that editing by transcript makes recording several takes and keeping the best one straightforward.