Google Adds Agentic Video Understanding to Selected Gemini Models
Google has made agentic video analysis available in the Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models through the Gemini API.

On September 1, 2026, Google announced the agentic video understanding feature for the Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite models. It is intended to analyze uploaded videos and YouTube videos through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
The new mode is designed to change how the model searches for information when working with video. Instead of processing content at a predetermined frame rate, it is supposed to dynamically search for and check relevant passages across visual frames, the audio track, and transcripts.
What agentic video understanding brings
With standard static video processing, the developer sets how much content or how many frames the model will evaluate. With long recordings, this increases the number of processed tokens and costs, while an important detail may fall outside the selected sample.
Google describes the new approach as an internal agentic process within the model. According to the specified question, Gemini is supposed to first search for potentially relevant segments and then verify them across the image, audio, and text transcript. The selection of passages thus moves from manual implementation in the developer’s application into the model itself.
The feature is activated by setting processing to “agentic” in the API configuration. Google says it will not charge a separate fee for this mode: standard Gemini API token prices are supposed to apply.
Available models and deployment
Google announced the agentic video analysis mode for three models in the Flash family:
- Gemini 3.7 Flash,
- Gemini 3.6 Flash,
- Gemini 3.5 Flash-Lite.
Developers can use it with videos uploaded to the system and with YouTube videos. According to the announcement, it is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. The announcement does not provide detailed feature limits or complete processing parameters.
Benchmarks currently come from Google
When introducing the feature, Google cited up to an 88% reduction in token usage and up to a 66% reduction in costs, along with an accuracy increase of up to 7%. These are, however, the manufacturer’s benchmark figures. No independent verification of the methodology or results was provided with the announcement, so it is unclear how these numbers will translate to different types of video and practical deployments.
The feature’s main significance is that video analysis in multimodal models is often token-intensive. If the stated results are also confirmed outside internal benchmarks, developers could search longer videos for specific events, audio information, or visual details at a lower cost.
What comes next
Google also announced plans to expand the feature to the Gemini app and the “Ask YouTube” service. The company has not yet confirmed the timing or scope of this availability beyond its own announcement.
As deployment continues, independent comparisons of accuracy, latency, and real-world costs with static video processing will be important to watch. Questions about API limits, documentation, and privacy protection when processing video content also remain open.
Sources
- Google Blog / Google DeepMind – The primary announcement from September 1, 2026 confirms the feature’s launch, supported models, availability channels, activation method, and the manufacturer’s benchmark claims.
Verified and updated: 09/01/2026 22:31



