arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

助けを求める声へ近づくロボットの音響ナビゲーション

AcousticDiffusion: Semantically Conditioned Audio-Guided Diffusion Policy for Search-and-Rescue Assistance

Iana Zhura, Didar Seyidov, Dmitrii Plotnikov, Hajira Amjad, Miguel Altamirano Cabrera and Dzmitry Tsetserukou

この論文をやさしく読む

ひとことで言うと

声がする方向だけでなく、助けを求める声かどうかも使って、ロボットが人へ近づく経路を作ります。音源の距離がはっきりしない状態も保持しながら移動します。

何に役立つ?

考えられる用途は、視界が遮られる場所で救助ロボットが人の声へ接近する支援です。実機で従来計画器より近くまで到達した結果を報告していますが、救助活動全体の成功を検証したものではありません。

この研究の面白いところ

音の意味の認識と、マイクで測る方向の不確実性を一つの経路生成に組み込んでいます。追加学習なしで四脚ロボットに適用し、録音音声を用いた検証とは別に実機の結果も示しています。

どこまで分かった?

合成検証の平均方位誤差11.20度に対し、実機では64.9度で、精度には大きな差があります。37.4%は比較対象に対する最終距離の改善であり、救助成功率ではありません。実機でも音源定位は不完全と明記されています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視界が悪い、または遮られた場所で活動する救助ロボットにとって、呼びかける人へ向かって移動することは重要な能力である。私たちは、人に向かう移動のための、意味情報を条件とする音声誘導型の拡散方策AcousticDiffusionを提案する。学習済みの音声認識器を固定して10.24秒のウィンドウを処理し、音声による選別と遭難状況を考慮した優先順位付けにより、認識出力を音源単位のナビゲーション上の役割へ変換する。 マイクロホンアレイの到来方向の測定値は、ロボットを中心としたベイズ的な鳥瞰信念場に逐次統合する。自己移動の補償により連続する観測を整列させ、方位測定に伴う距離の不確実性を保持しながら、音源位置を次第に絞り込む。意味情報を含む信念、最近の音響観測、音声特徴、ロボット状態を条件として、拡散モデルが経由点の軌道を生成する。 録音音声を用いた合成ナビゲーションの検証セットでは、AcousticDiffusionは終点の平均方位誤差11.20度を達成し、軌道の91.78%が呼びかける人の方向から30度以内に整列した。妨害音源の除外率は89.20~98.99%で、競合する話者がいる場合にHELPと指定された呼びかけ者を優先する割合は、ウィンドウの91.07%だった。追加の再学習なしにZSL-1四脚ロボットに載せてオンラインで動作させると、平均方位誤差は64.9度となった。比較対象のA*は98.2度、RRTは90.4度であり、提案手法の計画器の平均計算時間は6.07 msだった。 音響的な位置推定が不完全であるにもかかわらず、報告された最終的な音源までの平均距離は、ODAS(Open embedded Audition System)由来の誘導を使う従来の計画器の3.96 mから2.48 mへと減少し、37.4%改善した。これらの結果は、不確かな音響観測を、人の呼びかけにより近づく動作へ変換する本枠組みの能力を示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Navigating toward human callers is an important capability for rescue robots operating where visual contact is degraded or occluded. We present AcousticDiffusion, a semantically conditioned, audio-guided diffusion policy for human-directed navigation. A frozen pretrained audio recognizer processes 10.24 s windows, with speech gating and distress-aware prioritization converting recognition outputs into source-level navigation roles. Microphone-array direction-of-arrival measurements are recursively integrated into a robot-centric Bayesian bird's-eye-view belief field. Ego-motion compensation aligns successive observations, progressively constraining source position while preserving bearing-induced range uncertainty. The semantic belief, recent acoustic observations, audio features, and robot state condition a diffusion model that generates waypoint trajectories. On a synthetic-navigation validation set using recorded audio, AcousticDiffusion achieves a mean end-point bearing error of 11.20 degrees, with 91.78% of trajectories aligned within 30 degrees of the caller. Distractor rejection ranges from 89.20% to 98.99%, and the policy favors a HELP-designated caller over a competing speaker in 91.07% of windows. Deployed online on a ZSL-1 quadruped without additional retraining, it achieves a mean bearing error of 64.9 degrees, compared with 98.2 degrees for A* and 90.4 degrees for RRT, with a mean planner compute time of 6.07 ms. Despite imperfect acoustic localization, the reported mean final source distance is reduced from 3.96 m for the classical planners using ODAS-derived (Open embedded Audition System) guidance to 2.48 m, a 37.4% improvement. These results demonstrate the framework's ability to translate uncertain acoustic observations into closer approaches to human callers.

arXiv ID: 2609.21792 / 要約の誤りについて