arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

歩行者回避の性能を保ちながら移動ロボットを適応

Performance-Preserving Online Adaptation in Social Navigation via Diffusion Steering

Haruto Nagahisa, Kohei Matsumoto, Yuki Hyodo, Ryo Kurazume

この論文をやさしく読む

ひとことで言うと

人を避けながら目的地に向かう基本の移動能力を保ちつつ、その場所の通行慣習にロボットを合わせる方法です。基本の方策は固定し、生成時のノイズを選ぶ側を学習します。

何に役立つ?

考えられる用途は、場所ごとに人の動きや通行慣習が違う環境への適応です。既に学んだ移動性能を損なわずに調整することを目指しています。

この研究の面白いところ

複数の乱数シードで学んだ方策を統合した基盤に、ノイズ方策だけの学習を組み合わせています。環境適応と元の性能の維持を同時に評価します。

どこまで分かった?

実ロボットの評価はハードウェア・イン・ザ・ループ・シミュレーションです。実際の歩行者がいる配備現場での長期運用と同じではなく、要旨には安全保証や具体的な成功率はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人と共存する移動では、人間とロボットの複雑な相互作用をモデル化することが難しく、そのため深層強化学習が活発に研究されている。しかし、シミュレーションだけでは、多様な状況、ロボットの動力学、配備環境によって異なる社会的慣習を完全には再現できないため、配備先での微調整が有望である。その際、歩行者を避けて目的地に達するという移動の主目的を損なわないよう、基盤モデルの性能を保つ学習が必要になる。 本研究は、拡散方策を固定したままノイズ方策だけを学習する、強化学習による拡散ステアリング(DSRL)を適用し、性能を保つ学習を実現する方法を提案する。さらに、複数の乱数シードで学習した拡散ベースの強化学習方策を統合して基盤方策を構成し、学習性能を改善する。評価では、ほかの手法と比べ、提案法が性能を保ちながら効率的に学習できることを示す。また、社会的慣習への適応を通じた柔軟な行動制御と、ハードウェア・イン・ザ・ループ・シミュレーションを通じた実ロボットでの有効性を確認する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

In social navigation, modeling the complex interactions between humans and robots is difficult, and deep reinforcement learning has therefore been actively studied. However, because simulation alone cannot fully reproduce diverse scenarios, robot dynamics, and the social conventions that vary across deployment environments, fine-tuning in the deployment environment is promising. In doing so, learning that preserves the base model's performance is required, so as not to compromise the primary objective of navigation, namely avoiding pedestrians and reaching the destination. In this study, we propose a method that applies diffusion steering via reinforcement learning (DSRL), which trains only the noise policy while keeping the diffusion policy fixed, thereby achieving learning that preserves performance. Furthermore, we integrate diffusion-based RL policies trained with multiple seeds to construct the base policy, improving learning performance. Our evaluation shows that, compared with other methods, the proposed method enables efficient learning while preserving performance, and we confirm flexible behavior control through adaptation to social conventions, as well as its effectiveness on a physical robot through hardware-in-the-loop simulation.

arXiv ID: 2609.24317 / 要約の誤りについて