物体の種類が同じなら未知の形にも対応するロボット操作学習
KeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization
この論文をやさしく読む
ひとことで言うと
物体の形や大きさが変わっても使えるロボット操作方策を、3次元キーポイントから学習する研究です。
何に役立つ?
同じ種類の未知の物体を扱う操作方策の設計に役立つと考えられます。
この研究の面白いところ
点群から学んだ意味的な位置の対応を使い、操作軌道全体を予測しています。
どこまで分かった?
要旨は三つの操作課題と実世界での評価を述べていますが、具体的な成功率や実世界の試験条件は記載していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ロボット操作の汎化には、形、大きさ、姿勢が異なる未見の物体でも作業できる方策が必要である。しかし、従来の行動クローニング法は、個々の物体に固有の形状や外観に過度に適合し、新しい物体への転移が難しいことが多い。この研究はKeyGenを提案する。点群から標準化した意味的な3次元キーポイントを学習し、物体中心の構造化表現として方策学習に使う枠組みである。視覚運動拡散方策は、このキーポイントと物体中心の形状情報を条件に、操作の軌道全体を予測する。これにより、物体の個体間で一貫した幾何学的対応を得る。 物体カテゴリ内での汎化を評価するため、三つの操作課題を持つ写実的なシミュレーション・ベンチマークと、さまざまな物体について熟練者の軌道を作る計画ベースのデータ生成パイプラインを構築した。実験ではKeyGenは、姿勢が変わる条件の既知・未知の物体の双方で従来法を大きく上回った。物体当たりの実演数を増やした場合にも効果的に性能が向上し、物体の拡大縮小に対して頑健で、シミュレーションと実世界の操作の両方で高い性能を示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Generalization in robotic manipulation requires policies to perform tasks across diverse unseen object instances that vary in shape, size, and pose. However, conventional behavior cloning (BC) methods often overfit to instance-specific geometry and appearance, limiting transfer to novel objects. We introduce KeyGen, a framework that learns canonicalized semantic 3D keypoints from point clouds and uses them as structured object-centric representations for policy learning. A visuomotor diffusion policy conditions on these keypoints together with object-centric geometry to predict full manipulation trajectories, enabling consistent geometric correspondence across object instances. To evaluate category-level generalization, we construct a photorealistic simulation benchmark with three manipulation tasks and a planning-driven data generation pipeline that produces expert trajectories across diverse object instances. Experiments show that KeyGen significantly outperforms prior methods on both seen and unseen objects under pose variation, scales effectively with additional demonstrations per object, maintains robustness to object rescaling, and achieves strong performance in both simulation and real-world manipulation.
arXiv ID: 2609.28818 / 要約の誤りについて