arXiv論文メモ
新着一覧
cs.CV / cs.HC · 査読状況未確認

予算対応型ビデオ理解フレームワーク

Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding

Weitong Cai, Hang Zhang, Yukai Huang, Yiqiao Xie, Shan Gao, Jiankang Deng, Songcen Xu, Jifei Song, Zhensong Zhang

短い要約(全文訳を準備中)

長時間ビデオのエッジデバイスでの理解において、計算と帯域の制約下で時間構造を保持しつつ、視覚的属性を維持する方法を探る。言語記憶とピクセルの二重性を活用し、クエリごとに必要なフレームを取得する手法を提案。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-10(UTC)
最新改訂
2026-09-10 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text-only video memories lose fine-grained visual attributes. We observe a visual-textual duality: language memories carry long-range temporal structure better than dense frames, while pixels remain decisive for attribute-level perception. Building on this insight, we propose Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework. The edge runs a single offline captioning pass that builds a dual-track narrative index, an event-level story skeleton plus a clip-level micro-log, cached and reused across queries without re-captioning. At query time, a cloud-side MLLM reasons over the index in a story-first loop centered on a lightweight Visual-Need Router: a per-query gating module that triggers bounded keyframe retrieval only for perceptual questions (appearance, on-screen text, attribute disambiguation) and keeps temporal-structural questions in language space. The router turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length. Experiments on long-video benchmarks demonstrate strong accuracy-efficiency trade-offs while substantially reducing online visual processing.

著者のコメント

EMNLP 2026 Main Conference

arXiv ID: 2609.11899 / 要約の誤りについて