単眼RGB画像で視点の変化に強い多指ロボット操作
AnyViewDex: View-Invariant Dexterous Manipulation from RGB Observations
この論文をやさしく読む
ひとことで言うと
学習中に三次元の幾何情報を教え、実行時は単眼RGB画像と自己状態だけで視点変化に強い多指操作を行います。
何に役立つ?
深度センサーやカメラ較正への依存を減らし、ロボットの把持を異なる視点へ移すための学習方法です。
この研究の面白いところ
多視点の対照学習だけでは失われがちな位置情報を、シミュレーション中の絶対三次元座標の回帰で補います。学習時だけ使える情報と実行時の入力を分けます。
どこまで分かった?
実機で未見8物体・未較正6視点、480試行の把持成功率76.7%です。全アブレーション計2400試行とは区別され、任意の物体や複雑な操作全般での視点不変性を保証するものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
多指による器用な物体操作の視覚運動方策は、カメラ視点の変化に非常に敏感である。視点不変性を実現するため、最近の手法はRGB-Dや点群などの明示的な3次元情報への依存を強めているが、実環境への導入時に、ハードウェアへの依存、較正の必要性、センサーノイズへの弱さを生む可能性がある。本研究では、シミュレーション中に視覚表現へ幾何学的知識を組み込むことで、テスト時の明示的な3次元計測なしに視点不変な制御を実現できることを示す。 複数視点の対照的アラインメントと、学習時だけ利用できる3次元幾何の教師情報を組み合わせた非対称学習パイプラインAnyViewDexを提案する。シミュレーション学習中に物体の絶対3次元座標を回帰する補助目的によって、表現を幾何に結び付ける信号を与え、大域的にプーリングされた対照学習埋め込みにおける空間情報の崩壊を緩和する。導入時には、較正していない単眼RGBと自己受容感覚情報のみを用い、ゼロショットで方策を実行する。 この手法を、強化学習と生徒・教師間の蒸留の両方で検証する。16自由度のLEAP Handを取り付けたxArm7による実機評価では、未見の8物体と未較正の6視点にわたり、AnyViewDexの把持成功率は76.7%だった。試行数は480回、全アブレーション条件を合わせると2,400回である。この結果は、幾何に基づく単眼方策が、テスト時の深度情報なしにゼロショット転用できることを示している。 プロジェクトページ:https://anyviewdex.github.io/
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Visuomotor policies for multi-fingered dexterous manipulation are highly sensitive to camera viewpoint shifts. To achieve view invariance, recent methods increasingly rely on explicit 3D modalities like RGB-D or point clouds, which can introduce hardware dependencies, calibration requirements, and vulnerability to sensor noise during real-world deployment. In this work, we show that view-invariant control can be achieved without explicit test-time 3D sensing by encoding geometric knowledge into the visual representation during simulation. We present AnyViewDex, an asymmetric training pipeline that combines multi-view contrastive alignment with privileged 3D geometric supervision. By regressing absolute 3D object coordinates during simulated training, this auxiliary objective provides a geometric grounding signal that mitigates the spatial collapse of the globally pooled contrastive embedding. At deployment, the policy operates zero-shot using only uncalibrated monocular RGB and proprioception. We validate this approach across both reinforcement learning and student-teacher distillation. In hardware evaluation on an xArm7 with a 16-DoF LEAP Hand, AnyViewDex reaches 76.7% grasping success across eight unseen objects and six uncalibrated viewpoints (480 trials; 2,400 across all ablation conditions), indicating that geometrically grounded monocular policies transfer zero-shot without test-time depth. Project Page: https://anyviewdex.github.io/
arXiv ID: 2609.20107 / 要約の誤りについて