arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

大規模建物で平面図を使う自己位置推定を評価

SlugTrails: An Egocentric Benchmark for Floor Plan Localization in Large Buildings

Yunqian Cheng and Roberto Manduchi

この論文をやさしく読む

ひとことで言うと

大きな公共建物の中で、装着型カメラの映像と平面図から位置を求めるための評価データを提供しています。

何に役立つ?

住宅中心のデータでは分かりにくい、大規模建物での屋内測位の課題を評価できます。単一画像、静止しての複数視点、歩行中の連続観測を同じ建物で比較します。

この研究の面白いところ

三建物・六フロア、平面図外形22,089平方メートルのデータで、市販公開重みの性能はほぼゼロでした。追加学習で全ての学習可能な方式が改善し、別データLaMARへの転移も向上しています。

どこまで分かった?

建物データ不足が制約だとする証拠を示していますが、すべての屋内測位方式の限界を確定したわけではありません。位置・姿勢の許容誤差や観測形式ごとの評価値を区別する必要があります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

平面図に基づく屋内視覚自己位置推定は、専用インフラを設置せずに位置を求められる。しかし多くの手法は、実際の配備先である大規模な公共建物とは異なる、小さな住宅環境で開発・評価されている。本研究では、現実的な一人称センシング条件下の大規模屋内空間を対象とする、平面図自己位置推定ベンチマークSlugTrailsを導入する。大学構内の3建物・6フロアで記録した30 HzのAriaグラス映像、意味クラスと通行空間マスクを備えたCAD由来の平面図、レーザー測量した基準点を用いて平面図座標系に整列した軌跡を含み、平面図の外形で囲まれる面積は22,089 m²である。 1つの評価プロトコルで、限られた視野のもとで幾何情報を集める3つの実用的な方法、すなわち歩行中の単一フレーム、静止した状態での多視点スイープ、オドメトリ付き歩行映像列を扱う。異なる観測条件向けの手法を、同じ建物と正解情報で比較できる。代表的な幾何ベースおよび学習ベースの5システムを、それぞれ本来のセンシング構成で評価したところ、公式公開重みの成績はSlugTrails上でほぼゼロであり、歩行中の単一フレームにおけるR@1m30°は最大でも0.004だった。SlugTrailsで追加学習すると、学習可能なすべての手法群が3課題すべてで改善した。例えばF³Locは単一フレームで0.0から0.141へ、逐次設定で0.03から0.66へ向上し、観測が蓄積するほど改善が積み重なった。 同じ追加学習済み重みは、LaMARで学習していなくても、LaMARへのデータセット間汎化を改善する。逐次設定のR@1mは、F³Locで0.048から0.143、UnLocで0.063から0.127に向上した。一方、SlugTrailsだけでゼロから学習すると、公式重みから追加学習する場合を大きく下回る。これは、平面図自己位置推定の現在の制約がアーキテクチャよりも屋内データにあることを示す証拠である。データセット、プロトコル、ツールはhttps://github.com/Head-inthe-Cloud/SlugTrailsで公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Floor-plan-based indoor visual localization enables infrastructure-free positioning, but most methods are developed and evaluated in small residential environments unlike the large public buildings of real deployment. We introduce SlugTrails, a floor plan localization benchmark for large indoor spaces under realistic egocentric sensing: $30$ Hz Aria glasses recordings across three campus buildings and six floors ($22089$ m$^2$ of floor plan outline), CAD-derived floor plans with semantic classes and circulation space masks, and trajectories aligned into the floor plan frame using laser-surveyed anchors. One protocol covers three practical ways of gathering geometry under a limited field of view -- a single walking frame, a stationary multi-view sweep, and a walking stream with odometry -- so methods designed for different regimes are compared on the same buildings and ground truth. Evaluating five representative geometric and learned systems under their native sensing configurations, we find that stock checkpoints (official released weights) are near zero on SlugTrails (at most $0.004$ R@1m30$^{\circ}$ on walking single frames), while fine-tuning on SlugTrails improves every trainable family on all three tasks (e.g., F$^3$Loc $0.0 \rightarrow 0.141$ single-frame and $0.03 \rightarrow 0.66$ sequential), with gains compounding as observations accumulate. The same fine-tuned weights also improve cross-dataset generalization on LaMAR with no LaMAR training (sequential R@1m $0.048 \rightarrow 0.143$ for F$^3$Loc and $0.063 \rightarrow 0.127$ for UnLoc), whereas train-from-scratch on SlugTrails alone stays far below fine-tuning from stock weights -- evidence that floor plan localization is currently limited by indoor data rather than by architecture. We release the dataset, protocols, and tools at https://github.com/Head-inthe-Cloud/SlugTrails.

著者のコメント

8 pages, 4 figures. Preprint. Code and data: https://github.com/Head-inthe-Cloud/SlugTrails

arXiv ID: 2609.19876 / 要約の誤りについて