動画内の物体を時間と関係から探す検索グラフ
Structured Spatio-Temporal Evidence Graphs for Open-Vocabulary Object Retrieval in Videos
この論文をやさしく読む
ひとことで言うと
動画の中から「止まっている物」「ある場所にいる物」などを探すため、物体の動きと関係を索引に保存する。
何に役立つ?
時間や物体間の関係を含む自然言語の動画検索に役立つ。
この研究の面白いところ
フレーム単位の矩形では扱いにくい条件を追跡単位と物体の組のグラフで表し、三つの評価で改善した。
どこまで分かった?
持続する関係の検索は、物体の追跡が途切れないことと条件判定の調整に依存する。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
動画から自由な言葉で指定した物体を、検索時の費用を抑えながら探す必要がある。従来の索引は各フレームの領域を独立に保存し、画像と言語の類似度で検索するため、外見の指定には向く。しかし、停止している状態、画面内の場所、継続性、物体同士の関係のように、時間や関係を証拠とする条件には合わない。質問は物体の追跡単位や物体の組について述べるのに、索引は孤立した矩形だけを持つという不一致である。STEG-OVRは、継続して現れる物体を追跡単位のノード、時間的に対応する物体の組を関係の辺とする、時空間的な証拠グラフを提案する。外見、動き、画面領域での占有、相対的な幾何、速度の整合、記号的な関係を保存する。質問を物体、状態、場面、時間、関係の項目へ分け、必要な検索経路だけを起動した後、スコアを柔軟に統合し、一定の計算量で整合性を検証する。三種類の物体中心の診断実験で、平均適合率はBeachで0.0882から0.2073、Shibuyaで0.0834から0.1505、LOVO型の比較で0.701から0.743へ改善した。停止状態と場面内の領域指定で改善が大きい一方、持続する関係の検索は追跡の連続性と条件判定の調整に影響される。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Open-vocabulary object retrieval in videos requires answering free-form object queries under bounded query-time cost. Existing index-based systems typically store independent frame-level regions and retrieve them with vision-language similarity, which is effective for appearance queries but mismatched with predicates whose evidence is temporal or relational, such as stopped state, scene-region occupancy, persistence, and object interactions. We identify this gap as an evidence-unit mismatch: the query is expressed over tracklets or object tuples, while the index stores isolated boxes. To address it, we propose STEG-OVR, a structured spatio-temporal evidence graph for open-vocabulary object retrieval. STEG-OVR represents persistent objects as tracklet nodes and temporally compatible object pairs as relation edges, storing appearance, motion, scene occupancy, relative geometry, velocity compatibility, and symbolic relation evidence. A query is decomposed into entity, state, scene, temporal, and relation slots, which activate only the corresponding retrieval channels before soft score fusion and fixed-budget consistency verification. Diagnostic experiments on three object-centric settings show AP improvements from 0.0882 to 0.2073 on Beach, from 0.0834 to 0.1505 on Shibuya, and from 0.701 to 0.743 in a LOVO-style comparison. The gains are strongest for stopped-state and scene-region-occupancy queries, while sustained relations remain sensitive to tracking continuity and predicate calibration.
arXiv ID: 2609.23393 / 要約の誤りについて