SELF-EVOLUTION LONG-VIDEO TEMPORAL GROUNDING MULTIMODAL AGENTS

CoEvoWhen

Policy-Tool Coevolution for
Ultra-Long Video Temporal Grounding

Yiduo JiaMuzhi ZhuJinchuan ShiHao ZhongYuling XiKe LiuHao Chen†

Zhejiang University, State Key Lab of CAD & CG

† Corresponding author

+74.9%mIoU improvement
−11.4%Visual token cost reduction
5 benchmarks · 3 VLMsUltra-long video temporal grounding and QA

Qwen3.5-27B on ExtremeWhenBench, evolved skill vs. base skill.

Seeing Less yet Understanding More

A metal skimmer lifts fries from a fryer
50.6 MIN VIDEOCandidate Localization
and Motion Verification
View the trajectories
A couple walks along a tree-lined path in golden light
100.1 MIN VIDEOLong-Range Search
and Boundary Refinement
View the trajectories

Abstract

Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy–tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters.

During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition.

Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model.

Extensive experiments spanning five benchmarks and three VLMs show that policy–tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.

01 / COEVOLUTION

Policy–Tool Coevolution

Distilling ultra-long video grounding experience into a reusable external skill without updating model parameters.

02 / OBSERVATION

Coordinated Image–Video Observation

Enabling the VLM to acquire evidence autonomously without manually predefined coordination strategies or a stronger external planner.

03 / GAINS

Accuracy Gains and Cost Reduction

Improving grounding accuracy at lower visual token cost through separate evolution on different VLMs. The evolved skill also improves general long-video QA through direct transfer, without additional task-specific evolution.

CoEvoWhen Framework

A policy–tool coevolution framework that distills task experience into a reusable external skill for ultra-long video temporal grounding.

POLICIES

What evidence is needed, when to acquire it, and how to organize the search.

TOOLS

The form and granularity of the evidence the tools can present to the model.

Policy–Tool Coevolution

A frozen VLM executes ultra-long video temporal grounding tasks with an external skill comprising high-level policies and executable media tools. Across evolution rounds, an external skill updater distills transferable experience from execution trajectories and task feedback, refining policies for task planning and observation orchestration while synthesizing code to modify existing media tools or create new ones.

Policy-Guided Agentic Inference

Image-based observations can provide compact coverage of extended temporal ranges and facilitate comparisons across distant candidate regions, whereas video-based observations preserve local temporal continuity for reasoning about motion, event order, and temporal boundaries.

At inference, the VLM autonomously orchestrates media tools under the guidance of the evolved policy, using accumulated evidence to decide on subsequent observations without relying on a separate, stronger planning model.

Quantitative Results

Coevolution improves grounding accuracy while simultaneously reducing visual token cost. Consistent gains are also observed when skills are evolved separately on different VLMs. Directly applying the evolved skill to general long-video QA also improves performance without additional task-specific evolution.

Loading quantitative results…

Ablations & Analysis

Joint evolution achieves the best accuracy–cost combination.

Policy-only and tool-only variants both improve grounding accuracy over the base skill, confirming policies and tools as effective evolution targets.

More efficient evidence acquisition from ultra-long videos relies on both executable media tools to expand the available observation capabilities and high-level policies to select and orchestrate them.

↑ 5.0%higher IoU AUC
than policy-only
↓ 40.6%fewer visual tokens
than policy-only
↑ 16.7%higher IoU AUC
than tool-only
↓ 25.5%fewer visual tokens
than tool-only

Case Study

We analyze the execution trajectories of the base and evolved skills on two ExtremeWhenBench queries to illustrate how policy–tool coevolution changes evidence acquisition for action and scene localization in ultra-long videos.

Loading the recorded trajectories…

Citation

Open original ↗