arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

長時間動画理解のための自己改善型ツール設計

VideoResearcher: Self-Improving Tool Design for Long-Video Understanding

Dingqiang Ye, Dongdi Zhao, Kaishen Wang, Qingqiao Hu, Jingchen Sun, Yijun Liang, Yuqi Jia, Yiqiao Huang, Yunjie Tian, Jiaxing Zhang, Chuanyang Jin, Ke Zhang, Vishal M. Patel, Di Fu

この論文をやさしく読む

ひとことで言うと

長い動画を理解するエージェントが、足りない情報を見つけるための道具を自ら作り、試して改善する仕組みです。

何に役立つ?

動画向けの道具を人が何度も設計し直す負担を減らす用途が考えられます。モデルのパラメータを更新せず、利用する道具を改良します。

この研究の面白いところ

問題を解くループと道具を改善するループを分け、使用履歴から能力不足を診断して実行可能な道具を検証します。

どこまで分かった?

自己改善型エージェントの中で最高性能に達し、人の設計した上限へ近づいたと報告しています。要旨には具体的なベンチマーク値や開発コストの比較値はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

動画エージェントは長時間動画の理解で大きく進歩している。しかし効果的な動画エージェントシステムには、手作業による高コストで時間のかかる設計と試行錯誤が必要である。現在の自己改善手法は、影響の小さいプロンプトを改良するか、あらかじめ定義したマイクロツールを再結合するだけであるか、ハーネス最適化での収束に苦労している。この隔たりを埋めるため、VideoResearcherを提案する。これは、人間の研究者のように動画理解用の高影響ツールを自律的に設計、テスト、改良する、学習不要のマルチエージェント・フレームワークである。 VideoResearcherは、問題解決ループと進化ループという二重のループで動作する。ツール使用の軌跡を分析して能力の不足を特定し、専門化したエージェントを調整して実行可能なツールを開発・検証し、進化したツールを再利用して、その後の動画推論で証拠の取得を強化する。反復的なツール改良と検証を通して、モデルのパラメータを更新せずに証拠取得を段階的に強化する。VideoResearcherは自己改善エージェントの中で最先端の性能を達成し、人間が設計した上限にも近づいた。これは、コストの高い手作業のエンジニアリングを減らしながら、自律的なツール開発によってエージェントの能力を広げる、長時間動画理解の学習不要な方法を示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Video agents have made substantial progress in long-video understanding. Yet effective video-agent systems require costly, time-consuming manual design and trial and error. Current self-improvement methods either refine low-impact prompts, recombine predefined micro-tools, or struggle with convergence in harness optimization. To bridge this gap, we target high-impact video-tool with VideoResearcher, a training-free multi-agent framework that autonomously designs, tests, and refines tools for video understanding, like a human researcher. VideoResearcher operates through dual Solving and Evolving loops: it analyzes tool-use trajectories to identify capability gaps, coordinates specialized agents to develop and validate executable tools, and reuses evolved tools to strengthen evidence acquisition in subsequent video reasoning. Through iterative tool refinement and validation, it progressively strengthens evidence acquisition without updating model parameters. VideoResearcher achieves state-of-the-art performance among self-improving agents and approaches the human-designed upper bound, demonstrating a training-free paradigm for long-video understanding that expands agent capabilities through autonomous tool development while reducing costly manual engineering.

arXiv ID: 2609.19664 / 要約の誤りについて