編集できる3D建物を文章から作るProxyBuild
ProxyBuild: Text-Guided Structured 3D Building Generation with Mesh-Anchored Procedural Proxies
この論文をやさしく読む
ひとことで言うと
文章から3D建物を作る際、壁や部品を分けて後から編集できる構造を保つ方法。
何に役立つ?
仮想現実やデジタルツイン向けの、編集可能な建物モデル作成に役立つ可能性がある。
この研究の面白いところ
建物の外殻に部品を固定する中間表現を使い、形の解析と部品の配置を分けている。
どこまで分かった?
要旨は複数の指標での改善を述べるが、具体的な数値や実務での編集時間は示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
文章から3D建物を作る従来の生成モデルは、分離しにくい一つのメッシュや、操作できない描画結果を出すことが多い。手続き的なモデリングなら階層構造を持つ編集可能な建物を作れるが、規則を人手で書くのは大変で、言語モデルを使っても幾何の制約下で規則を解くのは難しい。ProxyBuildは構造を持つ建物生成のための混合的な枠組みである。新しい中間表現MAPPは建物の構成部品を幾何学的な外殻に固定し、生成を代理表現の予測と、代理表現から部品の実体化への二段階に分ける。まずMAPPの注釈を持つ建物データを作り、面と辺からなる二部グラフの符号器を学習させる。異種のメッシュグラフで位相的要素の特徴の関係を明示的に扱い、面と辺の意味上の役割を推定する。次に、言語モデルが読み取った文章の様式と属性値に基づき、空間配置の規則と厳格な制約を組み合わせて部品を検索・組み立てる。実験では、過度な平滑化、部品の衝突、構造の破綻を大きく減らし、さまざまな出所の意味付けのない外殻も正確に解析した。複数の指標で従来法を上回り、構造が明確で細部があり、後から編集できる3D建物を文章から頑健に作ったと報告する。仮想現実やデジタルツインで使う対話可能な資産の基盤となる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Text-guided 3D building generation holds tremendous application potential, yet existing generative models typically output inseparable single meshes or non-interactive rendered representations. While procedural modeling can generate editable buildings with hierarchical structures, rule authoring is laborious, and even with the aid of large language models (LLMs), it remains challenging to effectively solve procedural rules under geometric constraints. In this paper, we propose ProxyBuild, a novel hybrid framework for structured building generation. We introduce the Mesh-Anchored Procedural Proxy (MAPP) as a novel intermediate representation, which tightly anchors building components onto geometric shells, thereby decoupling the generation task into two phases: proxy prediction and proxy-to-asset instantiation. First, we construct a building dataset with MAPP annotations to train our designed face-edge bigraph encoder. By explicitly modeling the feature interactions of topological elements on heterogeneous mesh graphs, this encoder accurately infers the semantic roles of faces and edges. Subsequently, conditioned on textual styles and attribute parameters parsed by LLMs, we accomplish high-precision asset retrieval and assembly by integrating a spatial placement logic with hard constraints. Extensive experiments show that ProxyBuild not only significantly mitigates common issues in building generation such as over-smoothing, component collisions, and structural corruptions, but also accurately parses semantic-free shells from diverse sources. Outperforming prior baselines across various metrics, our method can robustly generate structurally clear, detail-rich, and post-editable 3D buildings from text, thereby providing a reliable and interactive content foundation for downstream applications such as virtual reality and digital twins.
著者のコメント
15 pages, 9 figures
arXiv ID: 2609.23386 / 要約の誤りについて