← All posts

How to Summarize a Video in ChatGPT or Claude Without Losing the Evidence

A five-part workflow for turning an authorized recording into a reviewable summary with source coverage, evidence references and human judgment.

You are the product lead who owes the team a decision-ready brief after a 45-minute customer call. The recording contains the buyer’s real problem, one objection, two possible next steps and a screen share that changes what “broken” means.

“Summarize this video” will produce something readable. That is not the same as producing something you can trust.

The better workflow is to make the assistant show its work: choose the authorized recording, define the summary’s job, ask what source coverage was actually available, require references for important claims, then review those claims against the video.

In ChatGPT, that workflow can use the Montage app in ChatGPT. In Claude, it can use the Montage connector in Claude. Montage supplies authorized recording evidence; ChatGPT or Claude composes the written summary. The distinction matters. You are reviewing an assistant-written interpretation of source evidence, not a native document exported by Montage.

The five-part summary brief

Before opening either assistant, write five lines:

  1. Source: Which recording may the assistant use?
  2. Reader: Who needs this summary?
  3. Decision: What should that reader understand or do next?
  4. Evidence: Which claims need a passage or time reference?
  5. Boundary: What must the assistant not infer?

For a customer call, the brief might be:

Use only the authorized recording titled “Onboarding review.” Write a one-page brief for the product team. Separate the customer’s stated problem, evidence shown on screen, decisions, open questions and next steps. Include a passage or time reference for every decision and action item. Do not invent an owner, deadline or level of urgency.

This gives the assistant a job. It also gives you rejection criteria.

Step 1: select the source deliberately

Connect the Montage app in ChatGPT or the Montage connector in Claude, then select the recording you are authorized to review. OpenAI says apps connect ChatGPT to services you already use so you can work with their information in a conversation; users can review requested services and permissions before completing authorization.[2] Anthropic says custom connectors connect Claude to workflow tools and data sources, and that OAuth authentication typically occurs during setup to grant specific permissions.[3]

Do not treat “connected” as “complete.” Ask the assistant to begin with a coverage statement:

Before summarizing, name the recording you accessed. State whether transcript, visual context and the requested time range were available. List anything you could not inspect.

If the answer shows that the recording is missing, still processing or only partially available, stop. A polished summary cannot repair absent evidence.

Montage gives the connected assistant the authorized recording context you selected; its public guide lists project, transcript-segment and Video Corpus access.[1] Montage does not certify, send or publish the resulting brief.

Step 2: define the summary’s job

Make the default output a one-page post-call brief for the product team: the customer’s stated problem, evidence, objections, decisions, action items and unresolved questions. Specify the length and reader. “Ten bullets for the product team” is more useful than “concise.”

Step 3: separate what was said from what the assistant inferred

Ask for separate sections:

  1. Directly supported: statements, decisions or actions present in the recording.
  2. Interpretation: the assistant’s synthesis of what those details mean.
  3. Unknown: details the source does not establish.

This prevents an ordinary inference from being mistaken for a decision. If someone says, “Maria can probably send it next week,” the summary should not silently become, “Owner: Maria. Due: Monday.”

Use this prompt:

Separate directly supported facts from interpretation. Quote no one unless you provide the exact passage for human verification. If an owner, date, amount or decision is not explicit, mark it unknown.

That caution is not ceremonial. Research on topic-focused dialogue summarization found significant factual errors in dialogue summaries across the models it evaluated.[6] Harvard’s AI guidance similarly tells users to review and correct AI output and warns that generated content may contain inaccuracies or fabricated facts.[4]

Step 4: make screen evidence part of the brief

A transcript can tell you that a speaker said, “The first screen is empty.” It may not tell you what was actually visible, which field was missing or whether the speaker clicked into a different state.

When the screen changes the meaning, ask for it:

Summarize the workflow shown on screen. For each key step, pair what the speaker said with the visible state that supports or contradicts it. Flag any visual detail you cannot determine confidently.

Montage’s current H0 product boundary supports spoken and visual recording context: it reads the screen, not just the transcript. Near-silent footage that depends entirely on visual interpretation remains outside the current promise, so use a recording with a transcript spine and review important visual claims yourself.

Step 5: require references where mistakes are expensive

Not every sentence needs a timecode. Decisions, commitments, figures, objections, quotations and claims about what appeared on screen do.

A practical review table is enough:

Summary claim Source reference Review result
Customer could not complete the first project Passage or time range Confirm / correct
The blank start caused the failure Spoken reasoning plus screen evidence Supported / inference
Team agreed to ship guided setup Exact decision passage Confirm / not decided
Maria owns the follow-up by Friday Exact owner and date passage Confirm / unknown

The table keeps the working summary honest without turning the final note into an audit log.

Worked example: the fluent summary that fails

For this guide, we used a rights-cleared 74.822-second fictional product-review recording. It contains five source ranges:

  1. 00:00.000–00:15.945 — the speaker discloses that the scenario is fictional and says weekly signups doubled.
  2. 00:15.945–00:30.290 — support conversations show that many teams never invited a colleague.
  3. 00:30.290–00:43.622 — the screen shows a blank first project with five choices and no example.
  4. 00:43.622–01:00.075 — the fictional team replaces the empty start with a guided first project and explicitly says no real metric improvement is being claimed.
  5. 01:00.075–01:14.822 — the speaker adds that experienced teams may still prefer a blank project.

A weak summary would be:

Signup growth hid an onboarding failure, so the team replaced the blank project with guided setup, which improved activation.

It is smooth and wrong. It removes the fictional-data disclosure, presents a fictional completed change as a real outcome, invents a measured activation result and drops the experienced-user exception.

A reviewable version is:

In a fictional example, signup growth did not show whether teams reached a shared workspace (00:00.000–00:30.290). The fictional team replaces a blank first project with a guided example (00:30.290–01:00.075) but explicitly reports no real metric result. It also notes that experienced teams may still prefer a blank start (01:00.075–01:14.822). The lesson to separate first-time and expert paths is supported; a real activation improvement is not.

The second version is not better because it is longer. It is better because every important statement has a source range and the uncertainty survives.

One prompt you can reuse

Use only the selected authorized call. Write a one-page brief with: stated problem, evidence, objections, decisions, action items and unresolved questions. Add a passage or time reference for every decision and action item. Do not infer owners, deadlines, sentiment or priority. Start with a coverage statement.

If the call contains a product walkthrough, add one line: pair each major spoken instruction with the visible state that supports or contradicts it, and flag any visual detail that cannot be verified.

What not to summarize this way

Use another process when:

  • you do not have permission to process the recording;
  • the source contains legal, medical, financial or employment decisions that require specialist review;
  • the important evidence is near-silent and purely visual;
  • exact quotations matter but you cannot inspect the source passage;
  • a missing recording would materially change the conclusion;
  • the output will be published without a named human reviewer.

Permission to process a private recording is not permission to reuse it publicly. Before turning a summary into a post, case study or article, remove or review names, organizations, dates, locations and distinctive details that can identify people.

The acceptance test

Before you share the summary, answer six questions:

  1. Does it name the source coverage?
  2. Does it separate direct support from interpretation?
  3. Are decisions, owners, dates, figures and quotations referenced?
  4. Did the review include important on-screen evidence?
  5. Are omissions and unknowns visible?
  6. Did a named human check the high-stakes claims against the recording?

If any answer is no, the work is not finished.

Select one authorized customer call. Ask what was available. Require references for the decision and follow-up. Check those claims before you send the brief.

Sources