AI can watch video. It does not watch yours when it decides who to cite.
Google's Gemini API takes a public YouTube URL and samples the video at one frame per second with audio. That is real video understanding, and anyone can demonstrate it in ten seconds. It is also close to irrelevant to whether your video gets cited, and the citation data shows why.
Gemini's share of YouTube citations in AI answers, the one engine that can process video (Otterly).
held by Perplexity, AI Overviews and AI Mode combined, none of which process video (Otterly).
the rate Gemini samples a YouTube video handed to it directly (Google).
One model really can watch a YouTube video.
Gemini accepts a public YouTube URL as a direct input. It samples visually at one frame per second, takes audio at 1Kbps, lets you address specific moments by MM:SS timestamp, and handles about an hour of video at default media resolution (Google, Gemini API docs).
That is not a transcript workaround. That is a model looking at frames.
It is also the model that cites YouTube least.
Otterly tracked more than 100 million AI citation instances over 30 days and broke the YouTube citations down by engine (Otterly).
| Engine | Share of YouTube citations | Can it process video? |
|---|---|---|
| Perplexity | 38.7% | No |
| Google AI Overviews | 36.6% | No |
| Google AI Mode | 19.6% | No |
| ChatGPT | 4.4% | No |
| Copilot | 0.5% | No |
| Gemini | 0.2% | Yes |
The one engine that can watch a video accounts for one citation in five hundred. The engines that cannot watch anything account for the rest.
Capability and retrieval are different systems.
Retrieval happens first. An engine assembling an answer selects candidate sources from an index built out of text, then summarizes them. If your video is not selected, its watchability never comes up.
A model processing roughly 300 tokens per second of video is not going to be run speculatively across every candidate source for every query when a description does the job for a fraction of the cost. That economics, not the technology, is what holds this answer in place.
Score your Video for AI Findability →
Your captions sit behind a door marked no crawlers.
On a YouTube watch page, caption data is served from YouTube's /api/ path, and YouTube's robots.txt disallows /api/ for general crawlers (youtube.com/robots.txt). A crawler that follows robots.txt is not permitted to fetch your captions.
That does not make transcript work pointless. Google crawls its own property under its own rules, and a corrected transcript improves everything downstream of it. It does mean the claim needs care. The retrieval path from YouTube is undocumented and partly blocked.
Which is why the transcript belongs on a page you own.
A watch page on your domain, with the transcript published as readable body text and VideoObject schema in the head, is the only version of your transcript you can be certain is retrievable. It is also the version where the citation credit lands on your domain instead of YouTube's.
What would have to change for this answer to change.
Two things, together. Video understanding would have to get cheap enough to run at retrieval scale, and the engines that dominate video citation today would have to adopt it. Watch one number: if Gemini's share of YouTube citations starts climbing off 0.2 percent, this page needs revisiting. Nothing else on this page is a leading indicator.
See where your videos stand.
Run your AI Findability audit and get your 0 to 100 score with the top gaps holding your videos back.
Score your Video for AI Findability →