Applied benchmark: Gemini 3.7 Flash agentic vs static video inspection

We ran a small matched benchmark to test the new agentic video inspection path against a static full-video pass while holding the model constant (Gemini 3.7 Flash).

Protocol

  • Six synthetic 10-minute videos with known brief events and editing decisions
  • Prompts and deterministic scoring frozen before the runs
  • No repair or retry
  • Capability, structured-output reliability, latency, token, and cost metrics

Five valid matched pairs

  • Brief-event recovery: 18/20 agentic vs 15/20 static
  • Edit-decision macro F1: 0.6807 vs 0.5481
  • Broad moment F1: 0.2667 vs 0.3000
  • Static used 26.42% fewer tokens and cost 23.01% less
  • One of six agentic outputs failed the required JSON contract

The conclusion is deliberately mixed: agentic inspection helped evidence-seeking and editing decisions, while static processing remained better on broad retrieval and efficiency. The sample is small and synthetic, with no human viewing panel, so this is exploratory—not a universal model ranking.

Protocol, raw outputs, scorer, and limitations: paperedits. com/ benchmarking/gemini-agentic-video-understanding-benchmark

Disclosure: I’m affiliated with PaperEdits, which published the benchmark.