Workflow templates

Summarize a YouTube video — or transcribe it, timestamp it, or describe what's on screen

Most summarizers read the caption track and hand you one paragraph. Sai asks what you actually need — plain transcript, timestamped transcript, key points, scene descriptions, or shareable clips — then keeps the video in memory so you can switch formats without starting over.

96
% success · 
453
 runs
youtube
youtube
The template
Copy prompt
Ask me for the YouTube URL and which analysis type I want — plain transcript, timestamped transcript, key points summary, scene and visual descriptions, or shareable clips with timestamps — then confirm before you start. Once I confirm: Fetch and process the video Extract the analysis type I chose, with timestamps where relevant Return the results in the chat and save them as a markdown file I can download Keep the full transcript in memory so I can ask for a different format afterwards without you re-fetching the video

See it run

The recording is a real session. The sheet on the right is what it produced.

Summarize a YouTube video — or transcribe it, timestamp it, or describe what's on screen
mp4

The run

Sai opens each profile, pulls the signal, and writes the row, live, in a real browser.

Summarize a YouTube video — or transcribe it, timestamp it, or describe what's on screen

The result

Eight columns, sorted by score, with a source link behind every claim.

Details

What you need

A YouTube URL. Sai asks which type of analysis you want before it starts.

What you get back

Your chosen output in the chat and as a downloadable markdown file — with timestamps you can click back to, and the transcript held in memory so you can ask for another format without re-processing.

How long it takes

A single pass covers the whole video — the run in the screenshot handled a 1 hour 25 minute talk. Asking for a different format afterwards is near-instant, since the transcript stays in memory.

Make it recurring

Point it at a channel or playlist you follow and run it weekly to keep notes on new uploads.

How do you summarize a YouTube video?

Paste the URL into a summarizer and you'll have a summary in under a minute. There are a dozen free ones, most don't require an account, and for a talking-head video they work well.

That's worth saying plainly, because the honest answer to "how do I summarize a YouTube video" is that this problem is largely solved and you shouldn't pay for it. Tools like NoteGPT and Recall handle long videos, produce chapters, and cost nothing.

So the useful question isn't how to get a summary. It's what to do when a summary isn't the thing you needed.

Why do most summarizers miss half the video?

Because they read the captions, and the captions only contain what was said.

Nearly every free summarizer works the same way: it pulls YouTube's caption track and sends that text to a language model. That's a sensible design and it's why they're fast and free. It also means the tool never sees a single frame.

For a podcast or an interview, that's fine — the audio is the content. For a lot of other video it isn't:

  • A software tutorial where the speaker says "then you click here." The caption records "then you click here." What they clicked is on screen and nowhere in the transcript.
  • A product demo where the interesting part is the interface, not the narration.
  • A design or code walkthrough where the screen changes constantly and the commentary is sparse.
  • A conference talk where the slides carry the data and the speaker says "as you can see."

"As you can see" is the phrase that exposes the whole approach. A caption-based tool cannot see. It transcribes the sentence and loses the thing the sentence was pointing at.

Sai can look at the video rather than only reading its captions, which is why scene and visual descriptions are one of the analysis types on offer. That's a structural difference, not a quality difference — it isn't that other tools describe the screen less well, it's that they have no access to it.

What can you ask for?

Five outputs, and it asks which one you want before it starts.

  • Plain transcript. Clean text, no timestamps. Best when you want to read or search rather than navigate.
  • Timestamped transcript. Every passage tagged, so you can jump back to any moment in the video.
  • Key points summary. Structured notes grouped by theme, with a timestamp on each point so a claim can be checked against the source.
  • Scene and visual descriptions. What's actually on screen — the thing caption-based tools can't reach.
  • Shareable clips with timestamps. Specific moments worth sending to someone, with the timestamp ready to paste.

The confirmation step before it runs matters more than it looks. Guessing wrong on a 90-minute video costs a full re-process, and asking one question up front is cheaper than being fast and wrong.

Can you change your mind after it runs?

Yes, and this is the part that behaves differently from a tool.

After processing, Sai keeps the full transcript in memory. Ask for the timestamped version after you got the summary, or for shareable clips after you got the transcript, and it produces them from what it already has rather than fetching the video again.

With a summarizer, switching output means running the whole thing over. Here the expensive step happens once and the formats are cheap afterwards. In practice you end up asking for two or three views of the same video, which is a workflow the one-shot tools quietly discourage.

What do you get to keep?

The output lands in the chat and as a markdown file you can download.

That distinction matters more than it sounds. A free summarizer gives you text in a browser panel — fine to read, annoying to keep. Markdown drops straight into Notion, Obsidian, a repo, or a doc with its formatting and timestamps intact.

It also makes the output an input. A timestamped summary of a competitor's launch video is the start of a research doc, and a set of clip timestamps is the start of a content plan.

When is a free summarizer the better choice?

When you want a paragraph about a talking-head video and nothing else.

If the video is an interview, the captions are good, and you just need the gist before deciding whether to watch — open a free tool. It'll take twenty seconds and cost nothing. Running a task for that is overkill and we'd rather say so.

This task earns its place when the video is visual, when you need a specific format rather than a generic summary, when you'll want more than one view of the same video, or when the output has to survive as a file in a workflow.

What can you build on top of it?

Video is an input to research, not a destination.

If you're working through a set of videos to understand a market or a competitor, our guide on using AI for market research and competitive analysis covers that end of it. If you're mining videos for content ideas, automated content creation picks up where the summary stops. And when you need the file rather than the text, we have a separate walkthrough on downloading YouTube videos to your computer.

Approach Reads captions Describes what's on screen Multiple formats from one fetch Downloadable file Best for
Free caption summarizers Yes
Fast and free
No
Never sees a frame
No
Re-run per output
Partly
Copy from the panel
A quick gist
Browser extension summarizers Yes No Partly
Preset views
No Summaries while browsing
Pasting a transcript into a chatbot Partly
If you fetch it yourself
No Yes
It stays in context
No
Copy and paste
Follow-up questions
Watching at 2x yourself Yes Yes
You're watching it
Yes Partly
Your own notes
Video you actually care about
Sai Yes Yes
Scene descriptions
Yes
Transcript held in memory
Yes
Markdown you own
Visual video and real workflows

Get more than a summary

Free your hands from the computer.

Run this Task