arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

長期計画と素早い動作を分ける両手操作モデルLiMA

LiMA: Bridging Long-term Imagination to Real-time Dexterous Manipulation via Asynchronous Diffusion

Ning Chen, Yankai Fu, Junkai Zhao, Qianpu Sun, Guocai Yao, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang

この論文をやさしく読む

ひとことで言うと

両手による物体操作で、長期の計画と素早い動作修正を別々の系に任せる方法である。

何に役立つ?

考えられる用途は、変化の速い接触を伴うロボット操作である。要旨では6種類の両手操作課題で評価した。

この研究の面白いところ

計画と実行を非同期に分け、比較手法より推論遅延を45.8%減らしながら成功率を測った。

どこまで分かった?

報告された成功率は評価した6課題での値である。他のロボットや環境での一般的な性能は要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

器用な物体操作には長期の見通しと素早い反応制御の両方が必要である。視覚・言語・行動(VLA)モデルは高度な推論に強い一方、物理的な動きや空間の細部を十分に捉えないことが多い。反対に、世界・行動モデル(WAM)は反復的な生成のため推論に時間がかかりやすい。これにより、モデルの意図が接触状態の素早い変化に追い付かない、時間的なずれが生じる。本研究はこれを克服するため、意図の計画と反応的な実行を系統的に分離した、非同期の二系統生成枠組みLiMAを提案する。 計算は複数の時間尺度からなる階層に分けられる。遅い系がまばらな長期の時空間的意図を生成し、速い系が密で高頻度な動きを細かく調整する。まばらな意図予測と密な行動軌跡を合わせるため、潜在Schrödinger Bridge Couplingを導入し、調整をエントロピー正則化した確率的輸送として定式化する。非同期に分けることで、Cosmos-Policyに比べて推論遅延を45.8%減らした。時間幅の異なる両手による器用な操作課題6種類で評価し、全体の成功率70.8%、部分課題の平均成功率78.9%を達成し、未見の状況でも性能を保った。プロジェクトのサイトは https://ccdcs.github.io/LiMA_repo/ で公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Dexterous manipulation demands long-term foresight and rapid reactive control. Vision-Language-Action (VLA) models, while proficient in high-level reasoning, often lack a fine-grained understanding of physical dynamics and spatial perception. Conversely, World-Action Models (WAMs) typically suffer from high inference latency due to iterative generation. These deficiencies result in a critical temporal misalignment where the model's intent fails to adapt to rapid physical contact changes. To overcome this fundamental bottleneck, we propose LiMA, an asynchronous dual-system generative framework that systematically decouples intent planning from reactive execution. LiMA organizes computation into a multi-scale hierarchy: a slow system handles sparse long-horizon spatiotemporal intent generation, while a fast system focuses on dense high-frequency motion refinement. To align sparse intent predictions with dense action trajectories, we introduce a Latent Schrödinger Bridge Coupling mechanism that formulates refinement as an entropy-regularized probabilistic transport process. LiMA reduces inference latency by 45.8% compared with Cosmos-Policy via asynchronous decoupling. Evaluated across six bimanual dexterous manipulation tasks spanning multiple horizons, LiMA achieves an overall success rate of 70.8% and an average subtask success rate of 78.9%, while maintaining performance in unseen scenarios. The project website is available at https://ccdcs.github.io/LiMA_repo/

arXiv ID: 2609.28431 / 要約の誤りについて