arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

三次元場面で部品と取っ手の関係から動きを推定

Geometric and Semantic Coupling for Interaction Understanding in 3D Scenes

Hanyang Kong, Xingyi Yang

この論文をやさしく読む

ひとことで言うと

三次元物体の部品と取っ手を結び付け、どこを操作するとどう動くかを推定する方法である。

何に役立つ?

考えられる用途は、ロボットが物体の操作可能な部分を理解すること。実証はArticulate3Dでの評価である。

この研究の面白いところ

動きの回帰器を学習せず、取っ手の位置を使う幾何デコーダーでヒンジを選んだ。

どこまで分かった?

評価値はArticulate3Dの検証条件でのもの。現実のロボット操作での成功率は要旨にない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

三次元場面で物体との操作を理解するには、動く部品、その動き、操作できる場所を一緒に記述する必要がある。本研究は、部品と取っ手の物理的関係を使ってこれらの出力を結び付けるSegment-Snapを提示する。学習済みの予測器が広い部品表面と小さな取っ手を特定する。幾何デコーダーは平面性と直立性の事前知識で動きを制約し、予測された取っ手の位置からヒンジ軸を選ぶ。動きの回帰器を訓練する必要はない。逆方向には、部品と取っ手を同時に予測する器が追加の取っ手候補を出し、その動きの種類を、それを含む部品から修正する。それぞれの情報の受け渡しは反復せず一回だけ適用する。 Articulate3Dの検証データでは、マスクと軸を固定した条件で、取っ手の情報により動きの条件を満たす平均適合率(AP)が13.74%から40.98%に上がった。追加の取っ手候補で取っ手APは24.63%から29.65%となり、部品による種類の修正でさらに0.98ポイント上がり、すべての文脈を使うと30.99%となった。訓練の反復、学習型デコーダーとの対照、対応する可視化によって、幾何と意味の証拠を組み合わせる利点と限界を調べた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Interaction understanding in 3D scenes requires a joint description of movable parts, their motion, and the regions through which they can be operated. We present Segment-Snap, which connects these outputs through the physical relationship between parts and handles. Learned predictors identify broad part surfaces and small handles. A geometric decoder uses planar and upright priors to constrain motion, then selects hinge lines using predicted handle locations, without training a motion regressor. Conversely, a joint part-and-handle predictor supplies additional handle candidates, whose motion classes are refined using containing parts. Each information transfer is applied once, without iterative feedback. On Articulate3D validation, handle guidance raises motion-gated AP from 13.74% to 40.98% at fixed masks and axes. Additional handle candidates raise handle AP from 24.63% to 29.65%; part-based class correction adds 0.98 points, and full context reaches 30.99%. Repeated training, learned-decoder controls and paired visualizations establish the benefits and limitations of combining geometric and semantic evidence for interaction understanding.

著者のコメント

Project page: https://hyokong.github.io/segment-snap-page/

arXiv ID: 2609.25247 / 要約の誤りについて