arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

画像内の指定点から物体の関節運動を推定する

Query-Conditioned Articulation Estimation from a Single Image

Abdelrhman Werby, Fabio Scaparro, Kai O. Arra

この論文をやさしく読む

ひとことで言うと

画像中で動かしたい部位の点を指定し、その周辺の関節の形や動きを推定する方法です。実寸への換算には、その点の奥行き測定を1回使います。

何に役立つ?

考えられる用途は、初めて見る可動物体をロボットが操作するための運動推定です。要旨では実機の移動マニピュレータによる操作試験も行っています。

この研究の面白いところ

画像だけでは実寸が決まらない問題を、指定点の奥行きを基準とした推定と、後から行うスケール復元に分けています。部位の領域分割に依存する既存の構成も見直しています。

どこまで分かった?

実機成功率70.2%は16部位・5視点分類・57試行の結果です。RGB画像だけで実寸を得たわけではなく、内部パラメータと指定点の深度測定を使います。全指標で優位だったという報告でもありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ロボットが関節を持つ物体の運動学的パラメータを推定できれば、さまざまな相互作用や操作が可能になる。推定にはロボットがその時点で観測できる情報を使う必要があり、多くの場合、それは初めて見る物体のRGB画像1枚にすぎない。現在の単一画像を用いる手法は、可動部の領域分割と関節パラメータの推定を結び付けているため、検出漏れや部位の対応付けの誤りに予測が左右される。また、単一視点ではスケールを除いてしか定まらない3次元形状を、実寸単位で回帰している。 本研究では、RGB画像1枚、2次元の問い合わせ点、カメラ内部パラメータから関節パラメータを推定するモデルQueryArtを提案する。QueryArtは、指定点を基準とし、その点の奥行きを単位とする3次元の関節幾何を推定するよう学習する。これによって、推定対象を画像だけから識別可能なものに保つ。その後、指定点で奥行きを1回測定すれば、スケールが与えられて実寸のパラメータを復元できる。選定した合成および実世界の関節データセットを組み合わせてQueryArtを学習する。複数のベンチマークで評価し、最近の比較手法と比較する。QueryArtは、分布外データを含め、関節に関する指標の大部分で最近の比較手法を上回る。実環境での能力を示すため、移動マニピュレータを用い、16の物体部位と5種類の視点にまたがる57回の操作試行を評価し、70.2%の成功率を達成する。コードと動画は https://abwerby.github.io/queryart/ で提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Enabling robots to estimate the kinematic parameters of articulated objects unlocks a wide range of capabilities for interaction and manipulation. The estimation has to happen from the information the robot currently observes, often just a single RGB image of an object it has never seen before. Current single-image approaches couple articulation part segmentation with articulation estimation, making their predictions vulnerable to missed detections and incorrect part associations, and they regress metric 3D geometry that a single view fixes only up to scale. We present QueryArt, a model that estimates articulation parameters from a single RGB image, a 2D query point, and camera intrinsics. QueryArt is trained to estimate the 3D articulation geometry relative to the queried point and in units of its depth, which keeps its target identifiable from the image alone. A single depth measurement at the query point then supplies the scale and recovers the metric parameters. We train QueryArt on a curated mixture of synthetic and real-world articulation datasets. We evaluate QueryArt on several benchmarks and compare it against recent baselines. QueryArt outperforms recent baselines on most articulation metrics, including on out-of-distribution data. To demonstrate the model's capabilities in real-world settings, we evaluate QueryArt on a mobile manipulator across 57 manipulation trials spanning 16 object parts and five viewpoint classes, achieving a 70.2% success rate. We provide code and videos at: https://abwerby.github.io/queryart/

著者のコメント

Code and video are available at https://abwerby.github.io/queryart/

arXiv ID: 2610.01726 / 要約の誤りについて