← All posts

What Should an AI Video Director Actually Let You Control?

A practical six-part control test for keeping source fidelity, editable output and human final judgment in an AI video workflow.

You are a content producer approving a short cut from a substantive recording. The opening is sharp. The captions look finished. It is ready for tomorrow’s launch.

Then you notice three problems: the cut starts after the speaker’s disclosure, the crop hides the screen evidence, and the final sentence removes the exception that made the advice honest.

The output looks edited. It was not directed.

The practical test for an AI video workflow is whether you can state the decisions that matter, inspect what the system did, and keep final judgment without rebuilding the video from scratch.

A useful director should give you control over six things.

1. Which source material is allowed

Start with source scope, not a creative prompt.

Write down:

  • the recordings the system may use;
  • the recordings it must ignore;
  • who owns or has permission to use them;
  • which speakers, sections or on-screen details are excluded;
  • whether outside stock, synthetic footage or generated imagery is allowed.

A simple direction might read:

Use only workshop-final.mp4. Ignore the rehearsal and audience Q&A. Keep the speaker’s disclosure in any cut that uses the fictional metric example.

That instruction defines evidence the system can use and one condition it cannot violate. “Make this engaging” does neither.

2. What the video must make clear

Name one viewer outcome and the beats required to earn it.

For a product lesson, the outcome could be:

After watching, the viewer should understand why signup growth can hide weak first value and know which two user paths to inspect separately.

Then list the required beats:

  1. the scenario is fictional;
  2. signups increased;
  3. support evidence contradicted the apparent success;
  4. the blank first project explained the failure;
  5. a guided first project was the proposed remedy;
  6. experienced users may still prefer a blank start.

This is the editorial contract. The system can choose pacing inside it. It cannot trade away the qualification that changes the lesson.

Brevity and completeness are different targets. If removing eight seconds makes the claim misleading, those eight seconds are part of the story.

3. How important decisions trace back to the source

A producer should be able to answer: why is this shot, sentence or caption here?

For spoken footage, ask for source timecodes or a transcript range. For screen recordings, name the visual frame or on-screen event that carries the evidence. For a jump between sections, record why the second section belongs.

Use a review table:

Output decisionSource evidenceAcceptance test
Open with the apparent signup win00:00.000–00:15.945Fictional disclosure remains audible or visible
Reverse the apparent win00:15.945–00:30.290“Arrival” is not presented as useful progress
Show the mechanism00:30.290–00:43.622Blank page, five choices and no example remain together
Present the remedy00:43.622–01:00.075No real metric improvement is implied
Preserve the exception01:00.075–01:14.822First-time and expert paths stay distinct

The viewer does not need a provenance log. The person approving the video needs enough source evidence to reject a misleading cut.

Treat this table as part of your directing brief and review process. It is a buying criterion, not a promise that every AI video product exposes decision-level provenance.

4. What the frame and captions are allowed to change

A vertical crop can remove the thing that proves the speaker’s point. Aspect ratio describes the relationship between frame width and height; changing that frame can change what remains visible.[4]

Direct the frame with evidence rules:

  • name the person or screen region that must remain visible;
  • state whether reframing may follow a speaker;
  • reject a crop that hides a demo, chart label or comparison;
  • choose a wider layout when the evidence cannot survive a tight crop.

Captions need a similar contract. W3C describes captions as the speech and non-speech audio information needed to understand the content.[2] Caption review should therefore cover:

  • transcript source;
  • speaker changes when they affect meaning;
  • product names, figures and proper nouns;
  • meaningful non-speech audio;
  • line breaks that do not imply a different sentence.

The format may adapt. The evidence must survive.

5. Whether the output stays editable

An AI-produced first cut is useful only if you can change your mind.

Before choosing a workflow, ask:

  • Can I reopen the result as an editable object?
  • Can I extend a source range when context is missing?
  • Can I remove a cut without losing every approved decision?
  • Can I correct captions and framing?
  • Can I compare the output with the source?
  • Can I reject one decision while keeping the rest?

Some of these controls may live in the product; others may require your own review sheet. Verify the exact workflow rather than assuming that “editable” means every decision is independently reversible.

Reusable direction can help when the same editorial judgment recurs. In Montage, Playbooks are new and rollout- and entitlement-gated. Treat them as an optional way to preserve approved direction, with case-by-case review still intact.

6. Who makes the final call

Human review should be a named stage with named rejection reasons.

NIST describes its AI Risk Management Framework as voluntary guidance for incorporating trustworthiness into the design, development, use and evaluation of AI systems.[3] Its Playbook says human-oversight processes should be defined, assessed and documented, and recommends documenting roles and responsibilities for that oversight.[5] For a video team, the lightweight version is an approval sheet:

  • Source fidelity: every material statement and visual comes from allowed footage.
  • Meaning: the cut preserves disclosures, qualifications and exceptions.
  • Frame: required visual evidence remains legible.
  • Captions: names, numbers and meaning-changing phrases are correct.
  • Brand: the result follows the approved voice and visual rules.
  • Delivery: duration, aspect ratio and destination match the brief.
  • Owner: one person approves or rejects the result.

Record the rejection reason: missing context, wrong source, hidden evidence, caption error, off-brand treatment or incorrect format. The next pass now has a concrete instruction.

Worked example: reject the attractive wrong cut

For this guide, we used an original 74.822-second fictional teaching recording. It contains no customer data and is cleared for Ukumi/Montage to reuse.

The recording moves through five beats:

  1. a fictional team celebrates signup growth;
  2. support conversations show that teams never reached a shared workspace;
  3. a blank first project with five choices reveals the mechanism;
  4. a guided first project supplies the proposed remedy and an explicit no-real-metric disclaimer;
  5. an exception separates first-time users from experienced teams.

The tempting short cut is 00:15.945–00:30.290. It contains a clean reversal: signup growth was real, but it did not prove useful progress.

We would reject it.

The cut drops the fictional disclosure, the screen evidence, the remedy and the exception. It is quotable, but incomplete.

A better direction brief is:

Use only this recording. Preserve the fictional-data disclosure, apparent win, contradictory support evidence, blank-page mechanism, remedy disclaimer and expert-path exception. Do not add outside footage or imply a measured result. Keep every evidence slide legible. Use the source transcript for captions. A human reviewer owns final approval.

The system may tighten silence and pacing. The required beats are fixed. That is enough freedom to produce a first cut without giving away the meaning.

Where Montage fits today

Montage is the AI production studio for directing videos from your own footage. Current product truth supports footage-grounded work: Montage reads the screen, not just the transcript; transcript-backed spoken intelligence supports spoken captions; Moment creation is bounded to under approximately three minutes per unit; outputs remain editable Montage objects; and final polish happens in Studio.[1]

A concrete test is to direct one high-risk Moment from authorized footage, then inspect and edit the resulting Montage object in Studio. Give it the required beats and the reason you would reject the cut. Check whether the disclosure, screen proof and qualification survive.

The workflow is Point → Direct → Produce:

  1. point Montage at authorized footage;
  2. direct the Moment you want and the constraints it must keep;
  3. review the editable result in Studio and make the final cut.

The source-traceability table and approval sheet remain the producer’s review method. This guide does not claim that Montage currently exposes every decision as a separate provenance record or lets you regenerate one decision independently.

Good fit

  • podcasts, interviews, events, courses and screen recordings with a spoken transcript spine;
  • Moments within the current under-approximately-three-minute unit boundary;
  • teams that want an editable first cut;
  • work where on-screen evidence matters alongside spoken words.

Outside the current promise

  • near-silent footage that depends entirely on visual interpretation;
  • autonomous full-length production;
  • automatic titles, sound effects, B-roll, publishing or scheduling;
  • guaranteed smart crops or thumbnails;
  • projects with no human available to review rights, meaning and final output.

The seven-question buying checklist

Before you adopt an AI video directing workflow, ask:

  1. Can I restrict it to the exact footage I authorize?
  2. Can I name the viewer outcome and the beats that must survive?
  3. Can I trace important output decisions back to transcript ranges or visual evidence?
  4. Can I protect required screen evidence during reframing?
  5. Can I correct captions, source ranges and cuts?
  6. Can I preserve reusable editorial rules without giving up case-by-case judgment?
  7. Can a named human approve or reject the result against explicit criteria?

If several answers are unclear, run one high-risk test before you commit: choose a Moment you would reject if it lost a disclosure, proof slide or qualification.

Bring that Moment, its required beats and its rejection criteria to montage.app. Direct one footage-grounded result, review the editable Montage object in Studio, and keep the taste and final cut.

Sources

[1] https://montage.app/llms.txt — Montage product facts [2] https://www.w3.org/WAI/media/av/captions — Captions/Subtitles — W3C WAI [3] https://www.nist.gov/itl/ai-risk-management-framework — NIST AI Risk Management Framework [4] https://helpx.adobe.com/premiere-pro/using/aspect-ratios.html — Adobe Premiere Pro — Aspect ratios [5] https://airc.nist.gov/airmf-resources/playbook/map — NIST AI RMF Playbook — Map

What Should an AI Video Director Actually Let You… | Montage Blog