arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

2次元の視覚言語モデルに3次元操作を行わせる

Task-Adaptive Grounded 3D-Programmers Using 2D VLMs

Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva, Jan-Nico Zaech, Luc Van Gool, Danda Pani Paudel

この論文をやさしく読む

ひとことで言うと

既存の2次元画像を扱うVLMに、共通座標系と作業結果へのフィードバックを与え、3次元の理解や編集を行わせる研究です。モデルの再訓練は不要とされています。

何に役立つ?

物体や室内シーンなどを、語彙を固定せずに指示して理解・操作・生成する用途が考えられます。既存VLMを利用する際の、軸や距離尺度の取り違えへの対処を提供します。

この研究の面白いところ

座標の表し方を整えるCCFと、課題に応じて結果を見直すTAFを組み合わせています。入力の表現と反復的な修正の両方を整えることで、3次元向けの再訓練なしに対応させる点が特徴です。

どこまで分かった?

要旨は多様な課題での実験結果を述べていますが、具体的な評価値やモデル別の比較、失敗例は示していません。全ての3次元操作を確実に行えるという保証までは読み取れません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

近年の視覚言語モデル(VLM)は優れた汎化能力と推論能力を示す一方、3次元理解はデータ規模、訓練の多様性、推論能力によって制限されている。本研究では、これらのモデルを単純に3次元へ拡張する代わりに、別の方法を取る。Canonical Coordinate Framing(CCF、標準座標系による枠付け)とTask-Adaptive Feedback(TAF、課題適応型フィードバック)という二つの新しい概念を通じ、3次元での位置付けと反復的なフィードバックループを導入して、強力な2次元VLMが3次元で確実に動作できるようにする。 CCFは、入力と出力の両方を共通のユークリッド座標系に固定する統一的な視覚表現として働き、軸の曖昧さ、尺度の不一致、基準の定まらない参照といった、3次元での位置付けに共通する課題を解決する。この構造化された3次元入力の枠付けを補うTAFは、課題に応じた動的なフィードバックで推論ループを閉じ、2次元VLMが本来の視覚的文脈の中で、多様なオープン語彙の課題を実行できるようにする。 この基盤の上に、CCFとTAFの能力を強力なVLMと組み合わせる、3次元の理解・推論・生成フレームワーク3D-Progを導入する。3D-Progは再訓練を一切必要とせず、物体単位とシーン単位の両方の課題で、オープン語彙による3次元の理解、操作、生成を行う。実験では、CCFとTAFの併用によって2次元VLMが幾何を認識する3次元プログラマーとなり、多様な3次元課題で、一貫性があり、解釈可能で高品質な結果を達成することを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operate reliably in 3D by introducing 3D grounding and iterative feedback loops with two novel concepts: Canonical Coordinate Framing (CCF) and Task-Adaptive Feedback (TAF). CCF serves as a unified visual representation that anchors both inputs and outputs to a shared Euclidean coordinate system, solving common challenges in 3D grounding such as axis ambiguity, inconsistent metric scale, and floating references. Complementary to this structured framing of the 3D inputs, TAF closes the reasoning loop with task-adaptive dynamic feedback that enables 2D VLMs to perform varied open-vocabulary tasks within their native visual context. Building on this foundation, we introduce 3D-Prog, a 3D understanding, reasoning, and generation framework that jointly employs the capabilities of CCF and TAF together with powerful VLMs. Without requiring any retraining, 3D-Prog performs open-vocabulary 3D understanding, manipulation, and generation across both object-level and scene-level tasks. Our experiments show that the joint use of CCF and TAF transforms 2D VLMs into geometry-aware 3D programmers, achieving consistent, interpretable, and high-quality results across diverse 3D tasks.

著者のコメント

18 pages, 9 figures, 11 tables

arXiv ID: 2610.02021 / 要約の誤りについて