arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

幾何制御と残差強化学習で本を狭い棚に差し込む

Hybrid Residual Reinforcement Learning for Contact-Rich Robotic Book Insertion

Tianyuan Liu, Rutherford Agbeshi Patamia, Benjamin Champion, Akansel Cosgun, Richard Dazeley

この論文をやさしく読む

ひとことで言うと

本を狭い棚へ差し込む最後の段階で、幾何学的な制御に学習した微調整を加えます。

何に役立つ?

わずかな位置ずれで引っ掛かる接触作業を、既知の動作構造を保ちながら改善するために役立ちます。

この研究の面白いところ

基本制御が挿入を担当し、PPOが有界な局所補正と手放すタイミングを選びます。実機では成功率が26.7%から63.3%へ改善しました。

どこまで分かった?

対象は把持と大域的な接近を終えた後です。シミュレーションの98.50%と実機の63.3%は別の結果で、極端に狭い隙間では局所補正の幾何学的限界が現れます。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

把持した本を狭い棚へ入れる操作は、小規模ながら難しい、多くの接触を伴う制御問題である。ミリメートル単位の姿勢誤差により、幾何学的には適切な接近でも、詰まり、解放の失敗、差し込み不足が起こり得る。本研究では、把持の獲得と大域的な接近の後に行う最終段階を対象に、既知の幾何情報と学習された挙動に制御をどう分担させるべきかを検討する。提案手法は、構造化された差し込みと所定位置への押し込みに公称のタスク空間コントローラを維持し、残差PPOによって範囲を限定した局所補正と解放時点の判断を行う。スクリプト化するのは、開く・後退する・再び閉じるという短い遷移だけである。 実機に用いる最終方策について、導入条件に合わせた512通りの固定条件でシミュレーション評価を実施した。独立した三回の学習での平均成功率は98.50%、標本標準偏差は0.23パーセントポイントであり、公称制御の成功率37.89%を上回った。実機xArm7でも、対応付けた30条件で計60試行を行い、同様の定性的な優位性を確認した。残差制御は成功率を26.7%から63.3%へ高め、失敗を22件から11件へ減らし、両コントローラの結果が異なった15条件のうち13条件で優れていた。 頑健性の試験では、初期化時の摂動を最大1.5倍にしても性能は87%を超えた一方、隙間が極端に狭い場合には局所補正の幾何学的限界が現れた。これらの結果は、幾何情報によって信頼できるタスク構造を維持し、固定規則では扱いにくい接触に敏感な挙動へ学習を集中させる、ハイブリッド設計を支持する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Placing a grasped book into a tight shelf is a compact but difficult contact-rich control problem: millimetre-scale pose error can turn a geometrically valid approach into jamming, failed release, or incomplete seating. We study this final phase after grasp acquisition and global approach, and ask how control authority should be divided between known geometry and learned behaviour. Our method retains a nominal task-space controller for structured insertion and seating, while residual PPO supplies bounded local corrections and decides when to release. Only the brief open-retreat-reclose transition is scripted. For the final policy used on hardware, a deployment-matched simulation evaluation over 512 fixed conditions yields 98.50 percent mean success (0.23 percentage-point sample SD) across three independent training runs, compared with 37.89 percent for nominal control. On the physical xArm7, 60 trials over 30 matched conditions show the same qualitative advantage: residual control raises success from 26.7 percent to 63.3 percent, reduces failures from 22 to 11, and wins 13 of the 15 matched conditions in which the two controllers differ. Robustness tests show that performance remains above 87 percent under initialization perturbations up to 1.5x, while very tight clearances expose the geometric limit of local correction. These results support a hybrid design in which geometry preserves reliable task structure and learning is concentrated on the contact-sensitive behaviour that fixed rules handle poorly.

著者のコメント

8 pages, 8figures

arXiv ID: 2609.19962 / 要約の誤りについて