音の方向を使うロボットの空間理解と移動
OmniEcho: Audio-Visual Spatial Understanding for Omni-Modal Embodied Agents
この論文をやさしく読む
ひとことで言うと
ロボットが音の方向と映像を合わせ、周囲の物の位置を考えて移動するためのデータとモデル。
何に役立つ?
音源を手掛かりにしたロボットの場面理解や移動の評価に役立つ。
この研究の面白いところ
実世界の空間音に加え、音源・映像・移動経路の位置関係が合った学習用音声を作っている。
どこまで分かった?
細かな音源位置の特定と距離推定には未解決の課題が残ると要旨が明記している。
v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人は音源の方向を捉えて視覚情報と合わせて考えられるが、身体を持つエージェントには難しい。特に、空間的な音の理解を実際の移動場面で評価・モデル化する方法は十分定まっていない。OmniEchoBenchは、空間的な音と映像の知覚、音・視覚・言語による移動を統一的に評価するベンチマークである。実世界の空間音・映像場面197件、質問と回答2,972組、30の実環境で集めた一次Ambisonics音声を伴う移動例900件を使い、六つの課題を含む。拡張しやすい教師データを作るため、音源、映像、エージェントの軌道の幾何学的な整合を保つ、制御可能な空間音のレンダリング手順も開発する。これを基にしたOmniEchoは、事前学習した意味的な音声処理経路に加えて、一次Ambisonicsの空間符号器を備える。実験では、空間的な音と映像の知覚で最高水準の性能となり、音に導かれる移動では従来の視覚・言語ナビゲーションに近い水準に達した。空間音が場面の推論と移動に有用な手掛かりとなる一方、細かな音源位置の特定と距離推定は重要な未解決課題である。
v2の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-23 · v2
- 査読・掲載
- 査読状況未確認
更新履歴
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Humans can effortlessly localize the direction of a sound source and integrate it with visual cues for reasoning, yet this remains challenging for embodied agents. In particular, it is still unclear how to effectively evaluate and model spatial audio understanding in embodied settings. To address this gap, we introduce \textbf{OmniEchoBench}, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation. OmniEchoBench comprises six tasks over 197 real-world spatial audio-visual scenes, 2,972 question-answer pairs, and 900 navigation samples with first-order ambisonics (FOA) audio collected from 30 real-world environments. To enable scalable training supervision, we develop a controllable rendering pipeline for spatial audio. It preserves geometric consistency among sound sources, visual observations, and agent trajectories. Building on this, we propose \textbf{OmniEcho}, a spatially aware omni-modal model. It introduces an FOA spatial encoder alongside a pretrained semantic audio pathway. Extensive experiments show that OmniEcho achieves state-of-the-art performance on spatial audio-visual perception. For our sound-guided navigation, OmniEcho reaches a performance level close to that of traditional vision-language navigation. These results demonstrate that spatial audio can serve as a valuable signal for embodied scene reasoning and navigation, while also highlighting fine-grained spatial localization and distance estimation as important open challenges. Our code and data will be available in https://github.com/PKU-VaLuE-Lab/OmniEcho/tree/main
arXiv ID: 2609.23407 / 要約の誤りについて