arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

大規模な意味地図を生成して物体探索を学習

NaviScale: Generating Large-Scale Semantic Map Datasets for Object Navigation

Chuanlin Lan, Yanwei Zheng, Weijian Liu, Zhitong Zhou, Xiao Zhang, Fuzhen Zhuang, and Dongxiao Yu

この論文をやさしく読む

ひとことで言うと

住宅の間取りと部屋の地図を組み合わせ、ロボットの物体探索を学ぶための大量の意味地図を作る研究。

何に役立つ?

考えられる用途は、実際の三次元環境を一件ずつ再構成せずにObjectNavの訓練データを増やすこと。

この研究の面白いところ

2万4000の間取りから19万2000の地図を作り、視野と遮蔽を考慮した部分観測も生成する。

どこまで分かった?

性能値は論文に記載された訓練・推論条件のHM3DとMP3Dでの結果。実機評価の詳細な数値は要旨にない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

実空間を動くロボットのナビゲーションには、未知の環境にも対応できる空間表現が必要だが、実際の三次元環境から大量の注釈付きデータを集めるのは難しい。本研究は、意味地図に基づいて目的の物体へ向かうObjectNavのためのNaviScaleを提案する。その予測器は、各訓練例について完全な三次元環境を復元せずに、部分的な意味地図と完全な意味地図の組から学習できる。 この枠組みは、実際の住宅の間取りと、MP3DおよびHM3DSemから抽出した部屋単位の意味地図・障害物地図を組み合わせ、大規模な訓練データを生成する。部屋間の拡大縮小で間取り全体の構造の多様性を増やし、部屋内の拡大縮小では、部屋の種類を合わせた異なる地図の組み合わせで同じ間取りを満たす。レイキャスティングによる可視性処理(VisRC)が、視野、センサーの届く範囲、遮蔽を考慮した部分観測へ地図を変換する。得られたデータセットには、1万2794物件に対応する2万4000の間取りから作った19万2000の意味地図が含まれる。 論文で述べる訓練・推論の設定で30万回の学習反復を行うと、予測器の構造を変えずに、HM3Dでは成功率64.3%、成功経路長による指標SPLが34.8%、MP3Dでは成功率43.1%、SPLが16.8%に達した。追加実験では、組み合わせた地図の品質、意味分割の誤りの影響、実機ロボットへの導入を評価した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Embodied navigation requires spatial representations that generalize across unseen environments, yet collecting large amounts of annotated data from real 3D environments is difficult. We propose NaviScale for semantic-map-based object navigation (ObjectNav), whose predictor can be trained on pairs of partial and complete semantic maps without reconstructing a complete 3D environment for every training sample. The framework generates large-scale semantic map training data by composing floorplans of real homes with room-level semantic and obstacle maps extracted from MP3D and HM3DSem. NaviScale increases data diversity in two ways: inter-room scaling increases floorplan-level structural diversity, while intra-room scaling fills each fixed floorplan with different combinations of room maps matched by room category. Visibility through Ray Casting (VisRC) converts the composed maps into partial observations that account for field of view, sensing range, and occlusion. The resulting dataset contains 192,000 semantic maps generated from 24,000 floorplans associated with 12,794 properties. With 300k training iterations and the training and inference settings described in this paper, the system reaches 64.3% SR and 34.8% SPL on HM3D, together with 43.1% SR and 16.8% SPL on MP3D, without changing the prediction architecture. Additional experiments evaluate the quality of the composed maps, the effects of semantic-segmentation errors, and deployment on a physical robot.

著者のコメント

14 pages, 8 figures; includes supplementary material

arXiv ID: 2609.27218 / 要約の誤りについて