画像の3次元推論で位置・方向・形状の表現を使い分ける
GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning
この論文をやさしく読む
ひとことで言うと
画像から空間を推論するAIで、位置・方向・形状を表す内部表現を分け、その表現を実際に回答に使わせる学習方法です。
何に役立つ?
考えられる用途は、2次元画像から方向や空間関係を答える視覚言語モデルの改善です。二つのベンチマークで性能を評価しています。
この研究の面白いところ
幾何表現が一方向に偏らないようにする工夫と、回答がその表現を経由するようにする工夫を組み合わせています。潜在表現の読み出しを遮断する比較により、その役割も調べています。
どこまで分かった?
方向の正答率が89.1%から25.8%へ変わる結果は、固定した128問での遮断実験です。二つのベンチマークのスコアとは別の評価であり、実環境のあらゆる3次元推論で同じ性能を示すとは限りません。要旨にそれ以上の限界の記載はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚言語モデルが進歩しても、2次元画像から3次元の空間を推論することは依然として難しい。テキストに基づく方法は、中間段階の幾何情報を離散トークンで記述するため、連続的な空間関係を忠実に表すことに限界がある。連続潜在表現はより豊かな表現を可能にするが、単一種類の潜在表現では、空間課題ごとに必要な手掛かりを明示的に分離できない。分解された空間潜在表現は、幾何学的な教師信号のもとで位置、方向、全体形状を別々に表すことで、この問題に対処する。それでも、幾何表現が一つの支配的な方向へ崩壊することがあり、注意機構を制限しない場合には、回答の学習中に潜在表現が十分に使われないことがある。 そこで、Common–Residual Geometry Alignment(CR-GEO)と経路を制御する最適化を組み合わせたGeoLatentを導入する。幾何状態を構造化しながら、潜在表現を介した回答学習を促す。CR-GEOは教師の幾何情報を共通成分と残差成分に分ける。経路を制御する最適化では、幾何と言語を同時に学習し、視覚情報に基づく回答学習を一時的に潜在表現経由へ向けた後、幾何学的な教師信号を保ちつつ、完全な注意を復元する。 条件をそろえた比較では、CR-GEOによって幾何表現の有効ランクが1.00から3.87へ上昇する。一方、ボトルネックで潜在表現の読み出しを遮断すると、固定した128問での方向の正答率は89.1%から25.8%へ低下する。復元後も、分化した幾何表現と潜在表現経由の視覚経路は、画像への直接アクセスと並んで利用可能なままとなる。GeoLatentはSPAR-Benchで73.0%、SPBenchで72.1%を達成し、両方で既報の手法を上回る。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and global geometry separately under geometric supervision. Yet the geometry representation can still collapse toward one dominant direction, and unrestricted attention can leave the latents underused during answer learning. We introduce GeoLatent, combining Common--Residual Geometry Alignment (CR-GEO) with routed optimization to structure the geometry states while promoting latent-mediated answer learning. CR-GEO separates shared from residual teacher geometry; routed optimization jointly trains geometry and language, temporarily directs visual answer learning through the latents, and restores full attention with geometry supervision. In controlled comparisons, CR-GEO raises geometry effective rank from 1.00 to 3.87, while blocking latent readout at the bottleneck lowers direction accuracy from 89.1% to 25.8% on 128 fixed questions. After recovery, the differentiated geometry representation and latent-mediated visual route remain available alongside direct image access. GeoLatent achieves 73.0% on SPAR-Bench and 72.1% on SPBench, outperforming previously reported methods on both.
著者のコメント
23 pages, 6 figures
arXiv ID: 2610.02091 / 要約の誤りについて