Tiếp nối đà phát triển thần tốc của phiên bản 3.7 Flash, Google đã chính…
Introducing agentic video understanding with Gemini
Today, we’re launching (agentic video understanding) across our latest models: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. This new capability improves accuracy while dramatically reducing token usage and costs for video analysis. Similar to agentic vision, which combines code execution with Gemini models’ native image understanding, agentic video understanding uses Gemini’s native video tools to improve performance and unlock new capabilities for video processing like sub-second moment retrieval, more accurate anomaly detection, precise counting and more.
The feature is available today for video uploads and YouTube videos via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
Benchmarks
Unlike current ‘static’ processing, where the model ingests the video at a fixed frames-per-second rate (default 1 FPS, adjustable via API), agentic video understanding pairs the model’s core reasoning with native video tools to dynamically search, scan, and inspect target video segments across visual frames, audio, and transcripts. Across standard video analysis benchmarks, Gemini models with agentic video understanding reduce analysis costs by up to 66% and token consumption by up to 88%, while improving accuracy by up to 7%.
These efficiency gains are especially pronounced on long-form video (from 10-minute how-to guides to 90-minute lectures and multi-hour recordings), where static processing forces developers to choose between high token costs or techniques that drop critical details.
Activating agentic video understanding drops token consumption by up to 88% and boosts accuracy by up to 7% with Gemini 3.7 Flash.
While these gains span all three supported models, Gemini 3.7 Flash with agentic understanding offers the best possible quality overall and the best combination of quality and cost efficiency, putting it at the accuracy-to-cost pareto frontier among tested models for video understanding.
Using agentic video understanding places Gemini 3.7 Flash at the accuracy-to-cost pareto frontier for video analysis.
How it works
Instead of static processing where the model ingests media streams at a fixed frame rate, agentic video understanding enables Gemini to take an active, goal-directed role in determining what to watch, at what speed, and through which modality (frames, audio, or transcript), fetching only the moments and signals needed. While developers could previously do this manually, with agentic video understanding, Gemini can accomplish it through an agentic loop, invoking an internal tool to load the relevant part of the video file, significantly reducing development overheads.
Capabilities and use cases
Agentic video understanding transforms how developers can process long-form video content across a variety of demanding applications.
- Sub-second moment retrieval: Pinpoint split-second state changes and tight cut boundaries that are easily missed at 1 FPS, making precise automated video editing possible.
- Long-form needle-in-a-haystack search: Answer complex queries across multi-hour videos without consuming millions of tokens.
- Anomaly detection: Resample interesting time windows at higher FPS to inspect rapid motion and subtle visual artifacts.
- Counting action & object: Accurately track repeated physical movements and distinct objects over time.
Real-world results
Many of our early access partners saw strong performance while testing with (agentic video understanding)Here’s what they have to say:
Revyl enables teams to automate the development and testing of iOS and Android applications using our cloud-based devices. While Revyl’s proprietary vision system can navigate apps like a real user, screenshot-based verification methods often miss fast-moving elements. By overcoming this limitation, our Agentic Video Understanding technology allows us to efficiently develop and test even those applications featuring complex visual effects.
Ethan Zhou, Growth Engineer, Revyl
“Our team put the Agentic Video Understanding model to the test alongside the existing Gemini system, using thousands of labeled videos and a real-time deepfake detection pipeline. While our models distinguish between authentic content and AI-generated material, Gemini translates those findings into easy-to-understand explanations. The new model delivered more accurate descriptions and context for long-form videos while reducing input tokens by 57%, enabling us to provide higher-quality results at a lower cost.”
Zohaib Ahmed, Co-founder & CEO, Resemble AI
Agentic Video Understanding made the most significant difference when handling long, complex raw footage—a critical factor for video-focused AI agents. In Mosaic’s real-world simulation benchmarks, this technology reduced average input tokens by 97%, nearly doubled the rate of adherence to constraints, and was the only mode capable of completing our 4.5-hour podcast processing case study.
Adish Jain, CEO, Mosaic
At Ponder, we built our own automated video understanding pipeline based on Gemini to identify valuable moments from raw footage. Google’s Agentic Video Understanding technology condensed that entire workflow into a single call, achieving performance comparable to our system while using approximately 3.5 times fewer input tokens.
Ibrahim Syed, Founding Engineer, Ponder
Getting started
Agent video understanding capabilities are available via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, deployed on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite versions. This service utilizes the standard Gemini API token-based pricing model and incurs no additional feature fees.
Google is also bringing performance and quality improvements in agentic video understanding technology to billions of users across its products. This feature will soon be rolled out to all Gemini app users, supporting both the Flash and Flash-Lite models. Additionally, in the coming months, this technology will also support the feature... “Ask YouTube” On YouTube video pages, leveraging the power of Gemini to provide higher-quality answers that closely align with the visual content of the video.




