arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

失敗した実演から避けるべき動作も学ぶロボット学習

CARF: Contrastive Attraction-Repulsion of Failure-Guided Flow Matching

Shuqi Zhao, Bang Du, Cheng-En Wu, Yichen Xie, Yixiao Wang, Masayoshi Tomizuka

この論文をやさしく読む

ひとことで言うと

失敗した実演を丸ごと捨てず、役立つ部分はまねし、失敗を招く部分からは遠ざかるように学ぶ方法です。

何に役立つ?

成功例だけでは不足するロボットの学習データを活用する用途が考えられます。曖昧な区間を除く仕組みも含みます。

この研究の面白いところ

良い部分の模倣だけでなく、悪い部分を明示的に避ける学習を同じ目的関数に組み込んでいます。評価器自体は成功実演とその摂動から学びます。

どこまで分かった?

シミュレーションと実環境の両方で改善を報告しますが、要旨には課題数や改善幅、評価器が誤判定する条件の記載はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ロボットの実演収集では、成功した実演だけでなく、不完全な軌跡や失敗した軌跡も得られる。既存手法は通常、失敗した軌跡からでも課題完了に向けて進んでいる区間を見つけて利用するが、失敗へ直接つながる決定的な行動をほとんど見落としている。本研究では、この二種類の区間は本質的に非対称な教師信号を与えると論じる。前進する区間は模倣すべきであり、失敗を決定づける区間は明示的に避けるべきである。 この観察に基づき、不完全なロボットデータから学ぶ、失敗情報に基づく対比的な引力・斥力の枠組みCARFを提案する。CARFは、成功した専門家の実演とそれに摂動を加えた結果だけで学習する、進捗に基づく重要度評価器を導入する。これにより、各ステップが課題完了にどれほど寄与するかを推定し、失敗軌跡の有益な領域を特定する。評価値は統一したフローマッチング目的を導き、方策を前進する行動へ引き寄せ、失敗を決定づける行動から遠ざけ、曖昧な区間は除外する。こうして不完全なデータをより包括的に利用しつつ、曖昧な失敗区間から信頼できない教師信号を取り込むことを避ける。シミュレーションと実環境での広範な実験は、多様な失敗状況で比較ベースラインを一貫して上回ることを示し、要素を取り除く実験も評価機構と引力・斥力機構の有効性を裏づける。研究サイト:https://zhao-sq.github.io/carf/#

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Robot demonstration collection often produces imperfect or failed trajectories in addition to successful demonstrations. Existing methods typically exploit failed trajectories by identifying segments that still make progress toward task completion, but largely overlook \textit{failure-critical behaviors} that directly lead to task failure. Here we argue that these two types of segments provide fundamentally asymmetric supervision: progressive segments should be imitated, whereas failure-critical segments should be explicitly avoided. Based on this observation, we propose CARF, a Contrastive Attraction-Repulsion of Failure-guided framework for learning from imperfect robot data. CARF introduces a progress-based importance scorer, trained solely on successful expert demonstrations and its perturbation results, to estimate step-wise contributions toward task completion and identify informative regions in failed trajectories. These scores guide a unified flow-matching objective that attracts the policy toward progressive behaviors and repels it from failure-critical ones, while excluding ambiguous segments. This enables more comprehensive utilization of imperfect data and avoids unreliable supervision from ambiguous failure segments. Extensive experiments in simulation and the real world demonstrate consistent improvements over competing baselines across diverse failure scenarios, with ablations further validating the effectiveness of the proposed scoring and attraction-repulsion mechanisms. Our website is https://zhao-sq.github.io/carf/#.

arXiv ID: 2609.21982 / 要約の誤りについて