Policy–Tool Coevolution
Distilling ultra-long video grounding experience into a reusable external skill without updating model parameters.
SELF-EVOLUTION LONG-VIDEO TEMPORAL GROUNDING MULTIMODAL AGENTS
Zhejiang University, State Key Lab of CAD & CG
† Corresponding author
Qwen3.5-27B on ExtremeWhenBench, evolved skill vs. base skill.
Seeing Less yet Understanding More
Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy–tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters.
During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition.
Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model.
Extensive experiments spanning five benchmarks and three VLMs show that policy–tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.
Distilling ultra-long video grounding experience into a reusable external skill without updating model parameters.
Enabling the VLM to acquire evidence autonomously without manually predefined coordination strategies or a stronger external planner.
Improving grounding accuracy at lower visual token cost through separate evolution on different VLMs. The evolved skill also improves general long-video QA through direct transfer, without additional task-specific evolution.
A policy–tool coevolution framework that distills task experience into a reusable external skill for ultra-long video temporal grounding.
What evidence is needed, when to acquire it, and how to organize the search.
The form and granularity of the evidence the tools can present to the model.
A frozen VLM executes ultra-long video temporal grounding tasks with an external skill comprising high-level policies and executable media tools. Across evolution rounds, an external skill updater distills transferable experience from execution trajectories and task feedback, refining policies for task planning and observation orchestration while synthesizing code to modify existing media tools or create new ones.
Image-based observations can provide compact coverage of extended temporal ranges and facilitate comparisons across distant candidate regions, whereas video-based observations preserve local temporal continuity for reasoning about motion, event order, and temporal boundaries.
At inference, the VLM autonomously orchestrates media tools under the guidance of the evolved policy, using accumulated evidence to decide on subsequent observations without relying on a separate, stronger planning model.
Coevolution improves grounding accuracy while simultaneously reducing visual token cost. Consistent gains are also observed when skills are evolved separately on different VLMs. Directly applying the evolved skill to general long-video QA also improves performance without additional task-specific evolution.
Loading quantitative results…
Policy-only and tool-only variants both improve grounding accuracy over the base skill, confirming policies and tools as effective evolution targets.
More efficient evidence acquisition from ultra-long videos relies on both executable media tools to expand the available observation capabilities and high-level policies to select and orchestrate them.
Both single-modality variants improve grounding performance through evolution, and image–video coordination yields further benefits: the evolved image+video skill outperforms both single-modality skills across all grounding metrics while using fewer visual tokens. These empirical results underscore that image-based and video-based observations are suited to different contexts in ultra-long video temporal grounding, and that strategically coordinating the two makes better use of their complementary strengths, achieving accuracy gains and cost reductions in tandem.
As evolution proceeds, the skill achieves higher grounding accuracy with lower visual token cost more consistently. Candidate omissions, boundary errors, and dynamic ambiguities exposed in trajectories drive coordinated adjustments to observation capabilities and their orchestration. Together, these advances suggest that policy–tool coevolution distills execution feedback into mutually adapted observation capabilities and policies, letting the VLM allocate its visual budget according to evidence needs for more accurate yet cheaper grounding.
We analyze the execution trajectories of the base and evolved skills on two ExtremeWhenBench queries to illustrate how policy–tool coevolution changes evidence acquisition for action and scene localization in ultra-long videos.