arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

複数センサーの点群を統合して人型ロボットを歩かせる

UniPoint: Unified Point-Level Sensor Fusion for Humanoid Locomotion Across Challenging Terrains

Sicen Li, Zhen Chu, Chao Li, Qiuguo Zhu and Jun Wu

この論文をやさしく読む

ひとことで言うと

LiDARと深度カメラの情報を1つの点群にまとめ、段差や隙間、細い障害物のある場所を人型ロボットが移動する方法です。

何に役立つ?

広い視野と局所的な障害物の把握を両立させる、機体搭載の知覚・移動制御の構成として参考になります。要旨では実機試行も報告されています。

この研究の面白いところ

画像をセンサーごとに処理するのではなく、早い段階で点群へ統合し、固定数のトークンにします。センサーの一部が失われたときの性能低下にも対応する設計です。

どこまで分かった?

学習対象は8地形ですが、実機評価は7地形・9設定で各20試行です。要旨には全体や条件別の成功率はなく、70cmの台や100cmの隙間という寸法だけから無条件の安全性を判断できません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

開かれた環境への導入では、人型ロボットが多様な地形を安全に横断する必要があり、知覚には広い観測範囲、局所的な正確さ、センサー故障への冗長性が同時に求められる。既存手法はこの3つを満たすのに苦労している。前方の深度カメラ1台や近傍の高さサンプリングでは範囲が狭すぎ、オドメトリで補正した標高地図は激しい動作でドリフトし、薄い垂直構造を見落とす。また、画像単位で符号化するとカメラ数とともに費用が増える。 本研究では、多源の点単位センサー融合に基づく、人型ロボットの全身移動の枠組みUniPointを提示する。360度LiDARと2台の深度カメラの測定を早期に融合し、機体座標系の1つの点集合にする。ボクセル化によって固定数のトークンへ再標本化し、線形自己注意と、固有受容感覚をクエリとする交差注意で符号化する。これにより順方向処理の費用をセンサー数から切り離す。点集合は立った薄い障害物を保持し、単一のセンサー様式が故障しても失われるのはトークンの一部だけなので、方策の性能低下は緩やかになる。 地形に応じた報酬、知覚劣化の注入、ドメインランダム化を使う1回の学習で、8種類すべての地形を扱う1つの方策を生成し、追加調整なしで機体搭載のRK3588へ導入する。DR02人型ロボットを使い、7種類の地形にまたがる9つの実環境設定で各20回の試行を行い、高さ70cmの台、100cmの隙間、薄い障害物、まばらな足場や狭い足場で方策を検証する。また、追加学習なしで屋外へも汎化する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Open-world deployment requires humanoid robots to cross highly heterogeneous terrain safely, with perception that simultaneously provides wide coverage, local accuracy, and redundancy against sensor failure. Existing approaches struggle to satisfy all three: one forward depth camera or nearby height sampling covers too little; odometry-corrected elevation maps drift under aggressive motion and miss thin vertical structures; image-level encoding costs grow with camera count. We present UniPoint, a humanoid whole-body locomotion framework built on multi-source point-level sensor fusion. Measurements from a 360{\deg} light detection and ranging (LiDAR) sensor and two depth cameras are early-fused into one base-frame point set. Voxelization resamples it to a fixed number of tokens encoded by linear self-attention and proprioception-queried cross-attention, decoupling forward cost from sensor count. The point set retains standing thin barriers; a single-modality failure removes only part of the tokens, so the policy degrades gracefully. A single training run with terrain-aware rewards, perception-degradation injection, and domain randomization produces one policy for all eight terrain types, deployed on an onboard RK3588 without fine-tuning. On a DR02 humanoid, 20 trials at each of nine real-world settings over seven terrain types validate the policy on 70-cm-high platforms, 100-cm gaps, thin barriers, and sparse or narrow footholds; it also generalizes zero-shot outdoors.

著者のコメント

8 pages, 9 figures, 6 tables. Submitted to IEEE Robotics and Automation Letters (RA-L). Video: https://youtu.be/Rd9YyfOxvmY

arXiv ID: 2609.23666 / 要約の誤りについて