両腕ロボットの行動と三次元の動きを同時に予測
JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation
この論文をやさしく読む
ひとことで言うと
両腕ロボットの動作と周囲の物体の将来の三次元的な動きを一緒に予測する方法。
何に役立つ?
片腕の動きがもう片腕の作業環境を変える、両腕協調の物体操作に役立つ。
この研究の面白いところ
シミュレーション16課題で平均成功率83.4%を達成し、実ロボットの三課題でも比較手法を上回った。
どこまで分かった?
実ロボットでの評価は三課題であり、すべての物体操作や環境での性能は要旨からは分からない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
両腕を協調させた物体操作は、片方の腕の動きが共有する三次元空間を変え、もう片方の腕にも影響するため難しい。しかし多くの拡散方策は、将来の幾何学的な影響を明示的にモデル化せずに行動を生成する。予測を使う方式でも、将来状態は補助的な教師信号か固定した条件として扱うことが多い。 本研究は、両腕の行動と、将来の三次元点の動きの軌跡を同時にノイズ除去する拡散方策JAMBを提案する。共有Transformer内で行動と軌跡の仮説を一緒に更新させることで、ノイズ除去の各段階で互いの情報を使って改善できる。さらに、複数の表現を共通の時空間座標系に置き、同時ノイズ除去中に幾何学を考慮した相互作用ができるようにした。RoboTwin 2.0の多様な両腕操作課題と実ロボットで評価し、行動だけを予測する方策や、異なる状態表現・学習目標を持つ将来予測法と比較した。 シミュレーションの16課題では平均成功率83.4%で、最も強い比較手法を23.9ポイント上回った。実環境の三課題では、行動のみの手法を50.0ポイント、幾何学の補助予測を使う手法を21.2ポイント上回った。評価した比較手法よりも、物が散らかった場面や学習分布から外れた背景への汎化も強かった。これらの結果は、両腕の協調操作で行動と物の動きを一緒にモデル化する枠組みの有効性を示す。プロジェクトのウェブサイトも公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Coordinated bimanual manipulation is challenging because the motion of either arm can alter the shared 3D scene and thereby affect the other arm. Yet most diffusion policies generate actions without explicitly modeling these future geometric consequences, while predictive variants typically use future state only as auxiliary supervision or fixed conditioning. We address this limitation by proposing JAMB, a diffusion policy that jointly denoises bimanual actions and future 3D point tracks. By allowing action and track hypotheses to evolve together within a shared Transformer, each can inform and refine the other throughout denoising. We further ground multimodal representations in a shared spatiotemporal coordinate system to facilitate geometry-aware interaction during joint denoising. We evaluate JAMB on diverse bimanual manipulation tasks in RoboTwin 2.0 and on a real-world robot, comparing it with action-only policies and alternative future-prediction approaches spanning different state representations and learning objectives. Across 16 simulation tasks, JAMB achieves an average success rate of 83.4%, outperforming the strongest baseline by 23.9 percentage points. On three real-world tasks, it outperforms the action-only and auxiliary geometry prediction methods by 50.0 and 21.2 percentage points, respectively. Beyond these performance gains, JAMB shows stronger generalization to cluttered scenes and out-of-distribution backgrounds than the evaluated baselines. Together, these results demonstrate the effectiveness of our joint action-motion modeling framework for coordinated bimanual manipulation. Our project website is available at https://jam-bimanual.github.io/
arXiv ID: 2609.25322 / 要約の誤りについて