We ran a small matched benchmark to test the new agentic video inspection path against a static full-video pass while holding the model constant (Gemini 3.7 Flash).
Protocol
- Six synthetic 10-minute videos with known brief events and editing decisions
- Prompts and deterministic scoring frozen before the runs
- No repair or retry
- Capability, structured-output reliability, latency, token, and cost metrics
Five valid matched pairs
- Brief-event recovery: 18/20 agentic vs 15/20 static
- Edit-decision macro F1: 0.6807 vs 0.5481
- Broad moment F1: 0.2667 vs 0.3000
- Static used 26.42% fewer tokens and cost 23.01% less
- One of six agentic outputs failed the required JSON contract
The conclusion is deliberately mixed: agentic inspection helped evidence-seeking and editing decisions, while static processing remained better on broad retrieval and efficiency. The sample is small and synthetic, with no human viewing panel, so this is exploratory—not a universal model ranking.
Protocol, raw outputs, scorer, and limitations: paperedits. com/ benchmarking/gemini-agentic-video-understanding-benchmark
Disclosure: I’m affiliated with PaperEdits, which published the benchmark.