arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

実行中のずれからロボットが復帰できるかを測る評価基盤

RoboRecover: Benchmarking Robot Policy Recovery under Execution Deviations

Yang Li, Chen Zhao, Zhuoran Wang, Jiankang Wang, Chao Shao, Yihan Lin, Haitao Shen, Jing Zhang

この論文をやさしく読む

ひとことで言うと

ロボットが作業途中で想定外の状態になったとき、元の作業へ戻れるかを測るベンチマーク。

何に役立つ?

初期状態からの成功率だけでは分からない、実行中の失敗からの復帰能力を比較するのに役立つ。

この研究の面白いところ

軌跡の途中まで行動を再生してずれた状態を作り、そこから元の課題を続けさせる。

どこまで分かった?

収録は RoboTwin と LIBERO の各1,000場面であり、実機環境全般への一般化は要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ロボット方策のベンチマークは多様な課題や事前設定された分布外の条件を扱うようになっているが、通常は所定の初期状態から始まる軌跡全体を評価する。こうした評価は初期場面と最終結果に注目し、動的な相互作用の過程にはあまり目を向けない。閉ループ実行中には行動や接触によって物体間の関係や課題の進行が変わり、復帰を要する想定外の中間状態が生じる。復帰には、課題の進み方がどう変わったかを推測し、関係を修正して元の目標を続ける必要がある。本研究は、実行中のずれからのロボット方策の復帰を評価するベンチマーク RoboRecover を導入する。軌跡から逸脱状態を選び、そこまでの行動列を再生して状態を再構成し、元の課題について方策を評価する。RoboTwin と LIBERO の両基盤に各1,000、合計2,000の場面があり、各基盤の訓練・試験分割は800対200に固定される。結果から、初期状態での性能だけでは復帰性能を決められず、方策の復帰能力は場面によって異なることが分かった。訓練分割は、復帰のための介入の研究にも利用できる。RoboRecover は、実行が引き起こした中間状態からの復帰を、ロボット方策の独立した評価軸として示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Robot-policy benchmarks increasingly cover diverse tasks and preset out-of-distribution conditions, but typically evaluate complete trajectories from predefined initial states. These evaluations often focus on the initialized scene and the final outcome, while paying less attention to the dynamic interaction process. During closed-loop execution, actions and contacts can alter object relations and task progress, producing off-nominal intermediate states that need recovery. Recovery requires a policy to infer how task progress has changed, correct the relevant relations, and continue the original goal. We introduce RoboRecover, a benchmark for robot policy recovery under execution deviations. RoboRecover selects deviation states from trajectories, reconstructs them by replaying action prefixes, and evaluates policies on the original task. RoboRecover contains 2,000 scenarios across RoboTwin and LIBERO, with 1,000 scenarios and a fixed 800/200 train/test split on each platform. Results show that initial-state performance does not determine recovery performance and policies exhibit different recovery strengths across scenarios. Using its training split, RoboRecover further supports study on recovery interventions. RoboRecover establishes recovery from execution-induced intermediate states as a distinct dimension of robot policy evaluation.

arXiv ID: 2609.28952 / 要約の誤りについて