arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

まとめて実行するロボット動作で誤差が縮むかを測る

Measuring the Stability Assumption Behind Action Chunking

Aryan Goyal

この論文をやさしく読む

ひとことで言うと

ロボットの動作に小さなずれを加え、動作列をそのまま続ける場合と再計画する場合に、ずれが自然に小さくなるかを測っています。

何に役立つ?

模倣学習したロボットが、予定から外れた後に戻れるかを評価する方法として役立ちます。回復動作を明示的に訓練する必要性を検討する材料です。

この研究の面白いところ

再計画すれば誤差が必ず縮むという見方を検証し、観測する時間幅によって誤差の増幅率の見積もりも変わると示しています。

どこまで分かった?

測定は三つのベンチマーク群の12課題に基づきます。摂動や分岐を網羅する訓練は結果から導いた提案であり、要旨ではその訓練による改善実証までは報告されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

行動をひとまとまりにするアクション・チャンキングは、行動クローニングで学んだ方策の性能を改善する。その理由として、時間的な一貫性、ホライズンの短縮、表現学習、誤差累積の軽減など、複数の仕組みが提案されてきた。本研究では代わりに、行動の誤差がシステムへ入った後に何が起きるかを調べる。 各状態で小さな行動誤差を加え、残りの動作列を再計画せず再生する開ループ実行と、摂動後に方策が再計画する閉ループ実行の二つについて、誤差がどれほど速く増大または縮小するかを測る。当てはめた変化率により、各状態を収縮、拡大、判定不能に分類する。三つのベンチマーク群の12の物体操作課題において、確信を持って安定と判定できる状態はまれで、伝播率を判定できる状態では誤差の増幅がよく起きることが分かった。さらに、測定される伝播率は当てはめる時間範囲に強く依存する。増幅は通常、初期に集中するため、短い観測窓は長い時間範囲の伝播を過大評価しうる。 最後に、これらのラベルを使って予測器を学習した。ある状態の開ループでの振る舞いはカメラ画像と自己受容感覚だけから推定できる一方、閉ループの伝播は、摂動後に方策がどう行動するかにも依存するため、一部しか推定できないことが分かった。これらの結果は、誤差累積の議論だけではアクション・チャンキングを十分に説明できないことを示唆する。受動的な開ループの動態も方策の再計画も、加えた誤差を一貫して収縮させず、再計画によって開ループの増幅が確信を持てる収縮へ変わることもまれである。このことは、閉ループの反応性が標準的な模倣学習から確実に生じると期待するのではなく、摂動と分岐木の網羅を重視した学習により、回復すべき逸脱を方策に経験させ、反応性を明示的に学習すべきだと示唆している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Action chunking improves the performance of policies learned by behavioural cloning, and several mechanisms have been proposed to explain why, including temporal consistency, horizon reduction, representation learning, and reduced error compounding. We instead study what happens to an action error once it enters the system. At each state, we inject a small action error and measure how fast it grows or shrinks under two execution regimes: open-loop, where the rest of the chunk is replayed without replanning, and closed-loop, where the policy replans after the perturbation. The fitted rate labels each state as contracting, expanding, or unresolved. Across twelve manipulation tasks from three benchmark suites, we find that confidently stable states are rare, while error amplification is common among states whose propagation rate can be resolved. We further find that the measured propagation rate depends strongly on the fitting horizon: amplification is typically front-loaded, so short windows can overestimate longer-horizon propagation. Finally, we train predictors on these labels and find that a state's open-loop regime can be recovered from camera frames and proprioception alone, while its closed-loop propagation is only partially recoverable because it also depends on how the policy acts after the perturbation. These results suggest that error-compounding arguments alone do not provide a complete account of action chunking: neither passive open-loop dynamics nor policy replanning consistently contracts an injected error, and replanning rarely turns open-loop amplification into confident contraction. This suggests that closed-loop reactivity should be trained explicitly, using perturbation- and tree-coverage-oriented training to expose policies to deviations they must recover from, rather than expected to emerge reliably from standard imitation learning.

著者のコメント

18 pages, 9 figures, 18 tables

arXiv ID: 2610.01626 / 要約の誤りについて