The Hallway Track
Engineering Insights

Agentic approaches to processing long videos with Gemini

Google Developers (Google I/O) · Sep 01, 2026 · Engineering Insights

Gemini's agentic video understanding selectively samples transcripts and frames to cut token costs while boosting accuracy.

“not only does it reduce the number of tokens it really needs, the performance increases because it can zoom in on certain functions and really pay attention to the things in the video that are actually important to the query”

Google's Gemini now supports an agentic loop for video analysis where the model decides which tools to call — get transcript, get frames, get audio — rather than ingesting the entire video at once. This selective approach keeps token counts manageable for long videos that would otherwise exceed 100k tokens. The technique is notable because performance improves alongside cost reduction, a rare double win in LLM video workflows.

gemini agentic-ai video-understanding multimodal token-efficiency google

Watch / read the original source →