arXiv論文メモ
新着一覧
cs.RO / cs.CV / cs.GR · 査読状況未確認

報酬プログラムを改良して人型ロボットに新課題を解かせる

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui

この論文をやさしく読む

ひとことで言うと

ロボットの制御器を再学習する代わりに、何をどの順序で達成したら報酬を与えるかというプログラムを改良します。既存の動作能力を、新しい移動・操作課題に組み合わせて使います。

何に役立つ?

学習済みロボットに新しい作業をさせる際、報酬設計と実行結果をつなぐ方法として役立つ可能性があります。多様な課題はシミュレーションで評価され、進化した技能の実機G1での自律実行も報告されています。

この研究の面白いところ

言語モデルが段階構成を直し、数値最適化器が定数を調整する分担です。検証済みのプログラムを技能として残すため、試行結果が次の計画に利用されます。

どこまで分かった?

シミュレーションでの多様な課題と実機での技能実行は区別する必要があります。要旨には実機で試した課題の範囲、成功率、必要な試行数はありません。再学習不要という説明は既存制御器の能力を利用するテスト時の処理についてです。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

私たちは、人型ロボットの移動を伴う物体操作におけるテスト時の進化を研究する。これは、制御器が学習したことのない課題を、既存の技能を転用し、自らの試行から改善し、学んだことを保持しながら、再学習なしで解くことである。中心となる洞察は、幅広い技能を備えた制御器には、新しい課題に必要な能力の多くが既に含まれているという点にある。その能力を引き出すには、計画と制御の間に、接触の多い多段階の相互作用を指定できるだけの表現力を持ち、同時に実行と測定が十分可能で、実行時のフィードバックによって経験に基づく計画を導けるインターフェースが必要となる。 InterEvolveは、2つの構成要素によってこのインターフェースを実現する。第一に、物体を考慮するforward-backward(FB)行動基盤モデルを開発する。固定した身体の事前モデルに物体に関する残差成分を加えることで、身体や物体に関する新しい報酬を、テスト時に移動・操作行動へ変換する。第二に、課題を報酬プログラムとして指定する。これは完了条件と調整可能な定数を持つ、段階的な報酬である。大規模言語モデル(LLM)エージェントは、実行時のフィードバックと検証済みプログラムの技能ライブラリーを参照しながら、文脈内でプログラムの構造を修正し、数値最適化器が定数を調整する。 各候補を並列のシミュレーション場面で検証することで、プログラムは、現在の課題に対して制御器の既存の運動能力を引き出し、転用し、組み合わせる新しい方法を探索し、反復とともに改善する。実験では、人が設計した報酬ではFBモデルの移動・操作能力の多くが引き出されていないのに対し、InterEvolveが進化させたプログラムは、ときには新しい戦略によって、その能力を引き出した。さらに、シミュレーションでは多様な課題、複雑な場面、長い時間にわたる技能の組み合わせに対応する行動を生成した。進化させた技能は、実機のUnitree G1でも、機体搭載の一人称視点の知覚を用いて自律的に実行される。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.

著者のコメント

Project page: https://sirui-xu.github.io/InterEvolve

arXiv ID: 2610.02196 / 要約の誤りについて