人型ロボットの学習済み動作を選んで忘れさせる方法
ForgetMimic: Motion Unlearning for Reinforcement Learning Humanoid Control
この論文をやさしく読む
ひとことで言うと
人型ロボットが覚えた複数の動作から、指定した動作だけを学習済み方策から除く方法を提案した。
何に役立つ?
考えられる用途は、不適切な動作や利用を取り消す必要がある動作を除き、残りの制御能力を保つこと。要旨では実機のUnitree G1とH2での実験を報告する。
この研究の面白いところ
対象動作の性能だけを下げ、ほかの動作の性能を維持することを目標にしている。学習内容の除去に失敗するロボット制御特有の機構も扱う。
どこまで分かった?
評価は要旨に記された2種類の人型ロボットと12動作について。法律上の権利への適合性そのものを検証した結果は要旨にない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人の動きを示すデータを利用した人型ロボットの制御では、強化学習によって多様で機敏かつ自然な移動動作が実現している。一方、物理的な人型ロボットの制御で高い性能が得られていても、学習済みの方策から特定の動作を除く方法は十分に研究されていない。この課題には、安全性とプライバシーに関する差し迫った動機がある。悪意のある動作、汚染された動作、性能の低い動作を除去することや、GDPRなどの規制における忘れられる権利の対象となり得る著作権保護された動作を除くことが重要である。そこで、現実世界で動く人型ロボット向けに設計した、動作単位の学習内容の除去法ForgetMimicを提案する。著者らはこれを同用途で初めての方法と位置付ける。N種類の動作で学習した方策πθについて、対象とするK種類の動作の性能を低下させつつ、残るN−K種類の動作の有効性を維持するのが中心的な考え方である。さらに、ロボット制御で学習内容の除去が失敗する原因となる2つの主要な学習機構を特定し、対処する。Dance、Fight、Flipなどを含む12種類の動作について、Unitree G1およびH2の人型ロボットで広範な実験を行った。結果は、指定した動作の記憶を効果的に除き、ほかのすべての動作の通常の動作を保てることを示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Humanoid control, leveraging human demonstrations, has achieved diverse, agile, and natural locomotion behaviors through reinforcement learning (RL). While this paradigm has yielded remarkable performance in physical humanoid control, how to eliminate specific motions from learned policies remains insufficiently explored. Addressing this issue is motivated by pressing safety and privacy concerns: the removal of malicious, poisoned, or suboptimal motions, as well as copyright-protected motions subject to the right to be forgotten under regulations such as the GDPR, is of critical importance. To this end, we propose {ForgetMimic}, the first motion-level unlearning method designed specifically for physical-world humanoid control. The core idea of ForgetMimic is as follows: given a policy $\pi_\theta$ trained on $N$ motions, our method degrades performance on a target subset of $K$ motions while preserving the effectiveness of the remaining $N-K$ motions. Furthermore, we identify and resolve two key training mechanisms in robot control that lead to unlearning failure. We conduct extensive experiments on the Unitree G1 and H2 humanoid robots across 12 motions, including Dance, Fight, Flip, and others. Experimental results demonstrate that ForgetMimic effectively eliminates memory of designated motions while maintaining the normal operation of all other motions.
著者のコメント
https://github.com/Zili1000/ForgetMimic
arXiv ID: 2609.28378 / 要約の誤りについて