深度画像の雑音に耐える四足ロボットの走破制御
DAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models
この論文をやさしく読む
ひとことで言うと
深度センサーに雑音があっても、手作業でフィルタを調整せずに四足ロボットが障害物を越えられるよう学習する方法です。
何に役立つ?
深度情報が不安定な環境で脚式ロボットを動かす際に役立つと考えられます。Unitree Go1による実機試験の障害物寸法が報告されています。
この研究の面白いところ
雑音付き入力からきれいな深度を再構成する学習と、雑音あり・なしの内部表現を近づける学習を組み合わせています。
どこまで分かった?
実証は要旨に記載されたUnitree Go1の障害物課題です。あらゆる雑音や地形での成功は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚に基づく脚式移動の手法は、学習時に雑音のない深度情報を仮定し、実機では手作業で調整した後処理フィルタに頼ることが多い。しかしフィルタのパラメータはほとんど公開されず、再現性を妨げるうえ、深度の雑音を放置すると性能は大きく低下する。学習手順に直接、雑音への頑健性を組み込めば、この依存をなくせる。自己受容感覚の入力ではそのような頑健性が検討されてきたが、脚式移動の深度知覚では同様の方法はほとんどない。本研究は、脚式移動のための雑音に強い知覚枠組みDAWNを提案する。世界モデルに二つの変更を加えて頑健性を組み込む。第一に、エンコーダーには雑音のある深度を入力し、再構成の目標には雑音のない深度を使うことで、モデルに入力の雑音除去を暗黙に学ばせる。第二に、対照学習によって雑音あり・なしの深度の潜在状態を整列させる。DAWNは特定の雑音モデルに依存せず、実機導入時に雑音分布を手作業で調整する必要がない。また、既存の世界モデル方式に比べて推論時の計算コストは増えない。手作業によるフィルタ調整をまったく使わず、学習した頑健な表現だけで、Unitree Go1の四足ロボットが生の深度観測からゼロショットでパルクール動作を行い、高さ最大18 cmの階段、幅最大70 cmの隙間、高さ最大45 cmの段差を通過した。要素除去実験では、雑音除去と対照的な整列がそれぞれ再構成と表現の水準で相補的に寄与し、組み合わせると効果が加算されることが示された。動画とコードが公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-based legged locomotion methods assume clean depth at training time and rely on hand-tuned post-processing filters at deployment. However, filter parameters are rarely disclosed, hindering reproducibility, and performance degrades substantially when depth noise is left unaddressed. Building noise robustness directly into the learning pipeline would eliminate this dependency. While such robustness has been explored for proprioceptive inputs, analogous approaches for depth perception remain largely absent in legged locomotion. We propose DAWN (Denoising and Alignment in World models for Noise-robustness), a noise-robust perception framework for legged locomotion, which builds noise robustness directly into a world model via two modifications: (1) feeding noisy depth to the encoder while keeping clean depth as the reconstruction target, forcing the model to implicitly denoise its input; and (2) applying contrastive learning to align the latent states of noisy and clean depth. Importantly, DAWN is not tied to a specific noise model, requiring no manual tuning to the noise distribution at deployment. Furthermore, it incurs no additional inference cost over existing world model-based methods. Without any manual filter calibration -- relying solely on the learned noise-robust representation -- DAWN achieves zero-shot quadruped parkour on a Unitree Go1: traversing stairs up to 18 cm, clearing gaps up to 70 cm, and mounting steps up to 45 cm from raw depth observations. Ablation studies show that denoising and contrastive alignment contribute at complementary levels -- reconstruction and representation, respectively -- and yield additive gains when combined. Videos and code are available at: https://dawn-parkour.github.io/
著者のコメント
8 pages, 6 figures. Accepted to IROS 2026
arXiv ID: 2609.29092 / 要約の誤りについて