Most summarizers read the caption track and hand you one paragraph. Sai asks what you actually need — plain transcript, timestamped transcript, key points, scene descriptions, or shareable clips — then keeps the video in memory so you can switch formats without starting over.
The recording is a real session. The sheet on the right is what it produced.
Sai opens each profile, pulls the signal, and writes the row, live, in a real browser.

Eight columns, sorted by score, with a source link behind every claim.
A YouTube URL. Sai asks which type of analysis you want before it starts.
Your chosen output in the chat and as a downloadable markdown file — with timestamps you can click back to, and the transcript held in memory so you can ask for another format without re-processing.
A single pass covers the whole video — the run in the screenshot handled a 1 hour 25 minute talk. Asking for a different format afterwards is near-instant, since the transcript stays in memory.
Point it at a channel or playlist you follow and run it weekly to keep notes on new uploads.
Paste the URL into a summarizer and you'll have a summary in under a minute. There are a dozen free ones, most don't require an account, and for a talking-head video they work well.
That's worth saying plainly, because the honest answer to "how do I summarize a YouTube video" is that this problem is largely solved and you shouldn't pay for it. Tools like NoteGPT and Recall handle long videos, produce chapters, and cost nothing.
So the useful question isn't how to get a summary. It's what to do when a summary isn't the thing you needed.
Because they read the captions, and the captions only contain what was said.
Nearly every free summarizer works the same way: it pulls YouTube's caption track and sends that text to a language model. That's a sensible design and it's why they're fast and free. It also means the tool never sees a single frame.
For a podcast or an interview, that's fine — the audio is the content. For a lot of other video it isn't:
"As you can see" is the phrase that exposes the whole approach. A caption-based tool cannot see. It transcribes the sentence and loses the thing the sentence was pointing at.
Sai can look at the video rather than only reading its captions, which is why scene and visual descriptions are one of the analysis types on offer. That's a structural difference, not a quality difference — it isn't that other tools describe the screen less well, it's that they have no access to it.
Five outputs, and it asks which one you want before it starts.
The confirmation step before it runs matters more than it looks. Guessing wrong on a 90-minute video costs a full re-process, and asking one question up front is cheaper than being fast and wrong.
Yes, and this is the part that behaves differently from a tool.
After processing, Sai keeps the full transcript in memory. Ask for the timestamped version after you got the summary, or for shareable clips after you got the transcript, and it produces them from what it already has rather than fetching the video again.
With a summarizer, switching output means running the whole thing over. Here the expensive step happens once and the formats are cheap afterwards. In practice you end up asking for two or three views of the same video, which is a workflow the one-shot tools quietly discourage.
The output lands in the chat and as a markdown file you can download.
That distinction matters more than it sounds. A free summarizer gives you text in a browser panel — fine to read, annoying to keep. Markdown drops straight into Notion, Obsidian, a repo, or a doc with its formatting and timestamps intact.
It also makes the output an input. A timestamped summary of a competitor's launch video is the start of a research doc, and a set of clip timestamps is the start of a content plan.
When you want a paragraph about a talking-head video and nothing else.
If the video is an interview, the captions are good, and you just need the gist before deciding whether to watch — open a free tool. It'll take twenty seconds and cost nothing. Running a task for that is overkill and we'd rather say so.
This task earns its place when the video is visual, when you need a specific format rather than a generic summary, when you'll want more than one view of the same video, or when the output has to survive as a file in a workflow.
Video is an input to research, not a destination.
If you're working through a set of videos to understand a market or a competitor, our guide on using AI for market research and competitive analysis covers that end of it. If you're mining videos for content ideas, automated content creation picks up where the summary stops. And when you need the file rather than the text, we have a separate walkthrough on downloading YouTube videos to your computer.