経験からロボットの長期作業手順を改良する適応的な仕組み
AdaHVLA: Adaptive Harnesses for Long-Horizon Vision-Language-Action Execution
この論文をやさしく読む
ひとことで言うと
ロボットの長い作業で、実行結果をもとに記憶と操作の連携手順を改良する仕組みです。
何に役立つ?
複数段階のロボット作業で、VLAモデルを管理する方針を改善する際に役立ちます。
この研究の面白いところ
失敗や成功の証拠、修正案、その効果をグラフで結び、次の試行に引き継ぎます。
どこまで分かった?
成功率の数値はシミュレーション結果です。実環境については安定した実行例が述べられていますが、同じ改善幅の数値は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚・言語・行動(VLA)モデルは、局所的な操作や指示への追従に優れる一方、持続的な記憶と計画が必要な長期課題では苦戦しやすい。タスクを管理する仕組みは、実行段階をまたいで履歴を保持し進捗を追うことで、エージェントの推論に持続的な文脈を与える。本研究では両者を組み合わせ、ロボットの経験を通じてコードによる連携方針を改良し、エージェントの推論・記憶をVLAの実行によりよく合わせる適応的な仕組みAdaHVLAを導入する。分離された複数エージェントによる適応過程では、証拠の分析、管理手順の修正、行動の評価を別々の作業文脈で行う。検証可能な連携についての仮説を修正の指針とし、その後の試行で予測した効果を評価する。状態を保持する改訂グラフが、実行時の証拠、仮説、修正、観測された効果を結び、代替となる管理手順や適応の記憶を残すことで、反復試行や異なる課題・環境での改善を導く。シミュレーションでは、NaVILA-LHのテストにおける平均成功率を22.5%から最大57.5%へ高め、三つのVLA基盤モデルにまたがる操作テストの成功率を初期の管理手順より最大30.8パーセントポイント改善した。実環境での運用も、適応した方針が複数の作業段階にわたり安定した実行を支える様子を示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-language-action (VLA) models offer strong local control and instruction following but often struggle with long-horizon tasks requiring persistent memory and planning. Task harnesses provide persistent context for agent reasoning by retaining task history and tracking progress across execution stages. To bring these complementary capabilities together, we introduce AdaHVLA, an adaptive harness that refines code-based coordination policies through robot experience to better align agent reasoning and memory with VLA execution. Its decoupled multiagent adaptation process separates evidence analysis, harness revision, and behavioral assessment into distinct working contexts, using testable coordination hypotheses to guide revisions and subsequent rollouts to assess their predicted effects. A stateful revision graph links execution evidence, hypotheses, revisions, and observed effects, preserving alternative harnesses and adaptation memory to guide refinement across repeated attempts and continued adaptation across tasks and environments. In simulation, AdaHVLA raises mean test success on NaVILA-LH from 22.5\% to as high as 57.5\% and improves manipulation test success across three VLA backbones by up to 30.8 percentage points over the initial harness. Real-world deployment further illustrates how the adapted policies support stable execution across task stages.
著者のコメント
9 pages, 5 figures
arXiv ID: 2609.29204 / 要約の誤りについて