人型ロボットの動作指令に接触力を組み込むOpt2VLA
Opt2VLA: Force-Aware Vision-Language-Action for Contact-Rich Humanoid Whole-Body Manipulation
この論文をやさしく読む
ひとことで言うと
人型ロボットへの指令に、どこへ動くかに加えて、どれくらいの力で接触するかを含める方法です。
何に役立つ?
物体に触れながら行う作業で、見た目の動きが同じでも必要な力が異なる場合の制御に役立ちます。言語に応じた力の調整をシミュレーションと実機で示しています。
この研究の面白いところ
VLAが力の目標を出し、全身制御器が追従する役割分担です。学習用の力とトルクを軌道最適化から与え、言語・視覚の判断を物理的な制御へつなぎます。
どこまで分かった?
評価は接触の多い3課題です。要旨には力誤差の具体値や長期運用の結果はありません。日常作業全般を人と同水準でこなせることを示したわけではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人型ロボットには日常環境で人と同水準の多様な作業を行うことが期待されており、その多くでは相互作用力を精密に調整する必要がある。近年の視覚・言語・行動(VLA)モデルは、意味に基づく計画や視覚運動制御で有望な結果を示している。しかし、既存の人型ロボットシステムは主に幾何学的な運動目標で行動を表し、運動追従に重点を置く全身制御器に依存しているため、相互作用力の明示的な推論や制御が限られている。この制約は接触の多い作業で特に問題になる。幾何学的には似た動作でも作業の文脈によって異なる力の状態が必要になり、接触後には視覚観測の信頼性が低下し得るためである。 本研究では、人型ロボットの全身操作においてVLAと制御の接続部へ明示的な力指令を導入する、力を考慮したVLAの枠組みOpt2VLAを提示する。単一の複数課題対応VLA方策が、幾何学的な運動目標と連続的な接触力の参照値を同時に予測し、課題ごとの強化学習(RL)に基づく全身制御器がそれらに追従する。規模を拡大しやすく物理的根拠のある教師信号を提供するため、明示的な力の参照値を備えた全身軌道最適化(TO)により、動力学的に実行可能で接触条件と整合する学習データを生成する。 接触の多い人型ロボットの3課題でOpt2VLAを評価し、力を明示的な条件として与えると、運動だけの制御より正確で一貫した力の調整が可能になることを示す。さらに、TOから得られる物理的根拠のあるトルク教師信号によって、力の追従精度と安定性が向上する。閉ループ評価では、シミュレーションと人型ロボット実機の両方で、言語に応じた力の調節を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Humanoid robots are expected to perform diverse human-level tasks in daily environments, many of which require precise regulation of interaction forces. While recent vision-language-action (VLA) models have shown promise for semantic planning and visuomotor control, existing humanoid systems primarily represent actions through geometric motion goals and rely on whole-body controllers focused on motion tracking, with limited explicit reasoning or control of interaction forces. This limitation is particularly relevant in contact-rich tasks, where geometrically similar motions may require different force regimes depending on the task context and where visual observations may become unreliable after contact. In this work, we present Opt2VLA, a force-aware VLA framework that introduces explicit force commands at the VLA-to-control interface for humanoid whole-body manipulation. A single multi-task VLA policy jointly predicts both geometric motion goals and continuous contact-force references, which are tracked by task-specific reinforcement learning (RL)-based whole-body controllers. To provide scalable and physically grounded supervision, we generate dynamically feasible and contact-consistent training data via whole-body trajectory optimization (TO) with explicit force references. We evaluate Opt2VLA on three contact-rich humanoid tasks and show that explicit force conditioning enables more accurate and consistent force regulation than motion-only control, while physically grounded torque supervision from TO further improves force tracking accuracy and stability. Closed-loop evaluations further demonstrate language-conditioned force modulation with Opt2VLA in simulation and on humanoid hardware.
arXiv ID: 2609.23968 / 要約の誤りについて