arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

汎用マルチモーダルモデルによる建物LOD3再構成

AstraLOD3: Zero-shot multimodal agentic reconstruction of LOD3 building models

Bryan G. Pantoja-Rosero

この論文をやさしく読む

ひとことで言うと

汎用マルチモーダルモデルとPython・Blenderを使い、建物のLOD3モデルを追加学習なしで再構成した。

何に役立つ?

建物の画像や点群から詳細な3Dモデルを作る処理の設計に役立つ可能性がある。

この研究の面白いところ

固定的な処理経路の代わりにエージェントが手順を選び、35回の実行で平均FRDS 0.9647を得た。

どこまで分かった?

評価は24棟のベンチマークを含む35回の実行である。適応的な修正や損傷対応は今後の課題として挙げられている。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

建物の詳細度LOD3での自動モデリングは通常、専用に設計した幾何処理や学習ベースの処理経路に依存し、建物の種類や入力資料の条件が異なる場合の柔軟性に制限がある。本研究は、汎用マルチモーダル基盤モデルAstraが、制限を設けた自律的な処理の枠組みで、追加学習なしにLOD3建物モデルを再構成できるかを調べる。AstraLOD3は、複数視点の画像、較正済みカメラ、フィルタ処理した疎なSfM点群を自然言語の再構成仕様と組み合わせ、AstraエージェントがPythonとBlenderを使って計算手順を動的に選択・実行する。ベンチマーク建物24棟を含む35回の実行で、平均FRDSは0.9647となり、幾何学的な一致度は従来の専用手法と同程度だった。条件を変えた比較実験では、再構成の指示、入力情報の種類、モデル設定、実行ごとのばらつきが結果に与える影響も明らかになった。これらの結果は、構造を持つLOD3再構成を固定的な処理経路ではなく、制約を設けたエージェントの処理として定式化できることを示す。今後は適応的な修正、利用者による修正、タスク固有の特化、損傷を考慮した再構成を調べる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Automated LOD3 building modeling typically relies on purpose-built geometric or learning-based pipelines, limiting flexibility across heterogeneous buildings and input evidence conditions. This study investigates whether Astra, a general-purpose multimodal foundation model, can address these limitations through zero-shot reconstruction of LOD3 building models within an agentic framework under bounded autonomy. AstraLOD3 combines multi-view images, calibrated cameras, and a filtered sparse SfM point cloud with a natural-language reconstruction specification, while the Astra agent dynamically selects and executes computational procedures using Python and Blender. Across 35 runs, including 24 benchmark buildings, AstraLOD3 achieved a mean FRDS of 0.9647 and geometric agreement comparable to that of previous purpose-built methods. Controlled ablations further revealed the effects of reconstruction guidance, evidence modalities, model configuration, and run-to-run variability. The results demonstrate that structured LOD3 reconstruction can be formulated as a constrained agentic process rather than as a fixed pipeline. Future work will investigate adaptive refinement, user-guided correction, task-specific specialization, and damage-aware reconstruction.

arXiv ID: 2609.28061 / 要約の誤りについて