arXiv論文メモ
新着一覧
cs.CV / cs.RO · 査読状況未確認

静止画像から物体の可動部と操作箇所を読み取るFunArt

FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents

Dennis Rotondi, Abdelrhman Werby, Kai O. Arras

この論文をやさしく読む

ひとことで言うと

動かしていない物体のRGB-D観測から、可動部分や取っ手と、その動き方を推定します。

何に役立つ?

ロボットが触れる前に、操作計画の初期情報を得る用途が考えられます。

この研究の面白いところ

生成3Dモデルの凍結済み圧縮表現を構造の事前知識として使い、機能部品と関節の性質を同時に復号します。

どこまで分かった?

Articulate3D上の評価で、機能要素は最強基準より6.7 AP50ポイント改善しています。実際に操作して成功した割合を示す指標ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人間の環境で効果的に動作するためには、ロボットは関節機構を持つ物体を識別し、その可動部や操作できる部分を分割し、運動学モデルを推定する必要がある。従来の関節機構を含むシーン表現は、通常、観測した相互作用から運動学を復元する。一方、静的なスキャンを用いる手法では、関節機構と機能的な操作要素を切り離して扱うことが多い。 本研究では、単一の静的な配置で撮影した、カメラ姿勢付きのRGB-D観測から、関節機構を考慮した機能的な3次元シーングラフを構築する枠組みFunArtを提示する。FunArtは個々の物体を再構成し、融合した形状をTRELLIS.2のO-Voxel表現に直接変換し、重みを固定した疎圧縮VAEを構造の事前知識として利用する。軽量なクエリベースのデコーダーは、コンパクトな物体単位の潜在表現と、表面に整合した密な特徴を組み合わせ、可動部と機能的な操作要素を同時に分割するとともに、運動の種類、軸、原点、範囲を推定する。 Articulate3Dデータセットでは、正解の物体入力を使う場合と使わない場合の両方で、可動部分割、関節機構推定、機能要素分割の各課題において最先端の性能を達成する。入力から出力まで一貫して処理する設定では、最も強いベースラインを、可動部でAP₅₀が1.5ポイント、原点と軸を同時に制約した条件でAP₅₀が2.8ポイント、機能要素でAP₅₀が6.7ポイント上回る。これらの結果は、生成的な3次元潜在表現が、物理的な相互作用の前にロボットの知覚と計画を初期化できる、行動に役立つ構造的手掛かりを符号化していることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

To operate effectively in human environments, robots must identify articulated objects, segment their movable and interactive parts, and estimate their kinematic models. Existing articulated scene representations typically recover kinematics from observed interactions, while methods operating on static scans often decouple articulation from functional interactive elements. We present FunArt, a framework that constructs articulation-aware functional 3D scene graphs from posed RGB-D observations captured in a single static configuration. FunArt reconstructs object instances, converts their fused geometry directly into the O-Voxel representation of TRELLIS.2, and exploits its frozen, sparse-compression VAE as a structural prior. A lightweight query-based decoder combines compact object-level latents with dense, surface-aligned features to jointly segment movable parts and functional interactive elements while estimating motion type, axis, origin, and range. On the Articulate3D dataset, FunArt achieves state-of-the-art performance across movable-part segmentation, articulation estimation, and functional-element segmentation, both with and without ground-truth object input. In the end-to-end setting, it outperforms the strongest baselines by 1.5 AP_{50} points for movable parts, 2.8 AP_{50} points under joint origin-and-axis constraints, and 6.7 AP_{50} points for functional elements. These results demonstrate that generative 3D latents encode actionable structural cues that can initialize robotic perception and planning before physical interaction.

arXiv ID: 2609.20673 / 要約の誤りについて