予算対応型ビデオ理解フレームワーク
Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding
短い要約(全文訳を準備中)
長時間ビデオのエッジデバイスでの理解において、計算と帯域の制約下で時間構造を保持しつつ、視覚的属性を維持する方法を探る。言語記憶とピクセルの二重性を活用し、クエリごとに必要なフレームを取得する手法を提案。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-10(UTC)
- 最新改訂
- 2026-09-10 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-10 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text-only video memories lose fine-grained visual attributes. We observe a visual-textual duality: language memories carry long-range temporal structure better than dense frames, while pixels remain decisive for attribute-level perception. Building on this insight, we propose Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework. The edge runs a single offline captioning pass that builds a dual-track narrative index, an event-level story skeleton plus a clip-level micro-log, cached and reused across queries without re-captioning. At query time, a cloud-side MLLM reasons over the index in a story-first loop centered on a lightweight Visual-Need Router: a per-query gating module that triggers bounded keyframe retrieval only for perceptual questions (appearance, on-screen text, attribute disambiguation) and keeps temporal-structural questions in language space. The router turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length. Experiments on long-video benchmarks demonstrate strong accuracy-efficiency trade-offs while substantially reducing online visual processing.
著者のコメント
EMNLP 2026 Main Conference
arXiv ID: 2609.11899 / 要約の誤りについて