動画内の出来事を探すエージェントの実行手順を自動改善
VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding
この論文をやさしく読む
ひとことで言うと
動画モデル自体の重みを変えず、その周囲の処理手順と指示を自動で改良して、出来事が起きた時間を探しやすくする枠組みです。
何に役立つ?
動画検索エージェントの失敗を踏まえた調整作業の自動化に役立つ可能性があります。要旨では複数ベンチマークでの位置特定性能の改善を報告しています。
この研究の面白いところ
段階間の接続は固定しつつ、内部の処理と指示には変更を許します。有望なコードの分岐を残して継続的に改善する設計です。
どこまで分かった?
指示改善の効果は一貫していた一方、コード変更の効果は評価条件によって異なります。具体的な改善幅や自動探索に必要な計算量は要旨には記載されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
動画の時間的グラウンディングは、自然言語の質問から動画内の出来事の位置を特定することを目的とする。重みを固定した動画言語モデルを中心に構築したエージェントでは、実行基盤が、質問から時間区間を予測する方法と、その予測を修正する方法を決める。この実行基盤を手動で改善するには、位置特定の失敗を診断し、エージェントの処理手順と指示の両方への変更を調整しなければならない。 本研究では、動画の時間的グラウンディング向けエージェント実行基盤を自動で進化させる枠組みVideoEvolveを導入する。段階間のインターフェースを保ちながら、処理手順と指示を進化の対象にできる穴埋め構造の実行基盤表現を用いる。分岐に基づく実行基盤の進化では、改善を継続できるよう有望なコードの分岐を保持し、実行時のフィードバックで局所的な編集を導き、検証によってどの改善を引き継ぐかを決める。 実験では、複数のベンチマークで位置特定性能が向上した。構成要素の分析では、指示の改善が一貫して性能向上につながる一方、進化したコードの効果は評価条件により異なることが分かった。これらの結果は、実行基盤の自動進化が、動画の時間的グラウンディングを改善する有効な方法であることを支持する。コードは https://github.com/bingjunluo/VideoEvolve で公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness determines how queries guide temporal predictions and how those predictions are refined. Manually refining these harnesses requires diagnosing grounding failures and coordinating changes to both agent workflows and instructions. We introduce VideoEvolve, a framework that automatically evolves agent harnesses for video temporal grounding. VideoEvolve uses a Cloze-Structured Harness Representation that preserves stage interfaces while leaving agent workflows and instructions open to evolution. Branch-Guided Harness Evolution preserves promising code branches for continued refinement, using execution feedback to guide local edits and validation to determine which improvements are carried forward. Experiments demonstrate improved grounding performance across multiple benchmarks. Component analyses identify instruction refinement as a consistent source of gains, while the benefits of evolved code vary across evaluation settings. Together, these results support automated harness evolution as an effective approach to improving video temporal grounding. Code is available at https://github.com/bingjunluo/VideoEvolve .
arXiv ID: 2610.01766 / 要約の誤りについて