arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

動画の空間推論の誤りを幾何演算で修正するCROSS

From Reasoning Failures to Composable Video Spatial Intelligence

Pengzhan Sun, Junbin Xiao, Ramanathan Rajaraman, Shiu-hong Kao, Angela Yao

この論文をやさしく読む

ひとことで言うと

動画内の位置や向きを答えるVLMの失敗を、認識不足、情報不足、測定の選択、座標や追跡の誤りに分けて調べています。幾何演算のライブラリで推論を補います。

何に役立つ?

動画の空間関係を回答するモデルやエージェントで、追加訓練せずに特定の誤りを減らす用途が考えられます。報告された平均得点はReVSIで4.3ポイント、DSI-BenchのSpatialClawで3.5ポイント上がっています。

この研究の面白いところ

問題全体の得点を見るだけでなく、予測と正解を共通の座標規約にそろえて原因を診断します。同じライブラリを、VLMへの文脈供給とエージェントのスキル呼び出しの両方に使えます。

どこまで分かった?

利用可能な証拠を使う仕組みであり、全ての知覚誤りを解消したとは述べていません。原要旨の評価部分には未展開の手法名記号が残っているため、訳では前述の手法を指す「本手法」としています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

空間推論ベンチマークは多様な課題で視覚言語モデルを評価するが、課題単位の得点だけでは、どの基礎能力が成功や失敗をもたらしたかは分からない。各課題では、空間的な証拠を取り出し、幾何を表現し、それに基づいて推論する必要がある。本研究では、共通のスキーマと座標の規約の下で、予測された空間的文脈と正解の空間的文脈を比較し、これらの能力を切り分ける。 この比較から、繰り返し生じる四つの誤りの原因が明らかになる。不正確な知覚、空間的文脈内の情報不足、誤った測定量の選択、参照座標系または位置・向きの追跡の誤りである。この診断に基づき、利用可能な証拠に対して働き、信頼できる動画空間推論を支援する、型付き幾何演算子と空間スキルの訓練不要なライブラリCROSSを開発する。このライブラリは、コードを書かないVLMへ検証済みの文脈を供給するか、SpatialClawエージェントへ呼び出し可能なスキルを供給する。 本手法を五つのベンチマークで評価する。本手法はReVSIの平均得点を55.9%から60.2%へ高め、DSI-BenchでのSpatialClawの結果を62.8%から66.3%へ改善する。これらの向上は、空間に関する規約を明示的に扱うことで、追加訓練なしに系統的な推論失敗を修正できることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities account for success or failure. Each task requires recovering spatial evidence, representing geometry, and reasoning over it. We disentangle these capabilities by comparing predicted and ground-truth spatial context under a shared schema and coordinate contract. This comparison reveals four recurring sources of error: inaccurate perception, missing information in the spatial context, selection of the wrong measurement, and errors in reference frames or in tracking position and orientation. Guided by this diagnosis, we develop CROSS, a training-free library of typed geometric operators and spatial skills that function over available evidence to support reliable video spatial reasoning. The resulting library supplies verified context to non-coding VLMs or callable skills to a SpatialClaw agent. We evaluate \methodname{} on five benchmarks. \methodname{} raises the average score from 55.9\% to 60.2\% on ReVSI and improves the SpatialClaw result from 62.8\% to 66.3\% on DSI-Bench. These gains demonstrate that explicit handling of spatial conventions can repair systematic reasoning failures without additional training.

arXiv ID: 2610.01999 / 要約の誤りについて