断面を重ねた表現で立体のつながりを保つ3D生成
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
この論文をやさしく読む
ひとことで言うと
立体を大量の小さな箱で表す代わりに、重なり合う断面で表す3D生成法です。細い部分や離れた場所同士のつながりを保ちつつ、計算負担を減らします。
何に役立つ?
高解像度の立体生成で、形のつながりとメモリ・時間の両方を評価するのに役立ちます。実験では学習時メモリと推論時間の削減を報告しており、製品制作への適用は考えられる用途です。
この研究の面白いところ
断面を圧縮表現に使うだけでなく、穴や連結成分などを捉える位相情報を隣り合う断面間で合わせます。複数方向の断面を共通の3D空間で協調させる構成です。
どこまで分かった?
数値によって比較相手が異なり、精度は最も強力なベースライン、トークン数は次にコンパクトな方式や疎・階層型方式との比較です。要旨にはデータセット、ハードウェア、各指標の詳細な測定条件は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
高解像度の3D生成では、ボクセルの潜在表現と、まず有効な構造を予測し、その後に局所形状を合成する多段階パイプラインへの依存が強まっている。この設計は有効だが、連続した表面を多数の局所トークンに分断し、生成コストを増やし、細い形状やつながりの多い形状では位相的な整合性を損なうことが多い。私たちは、コンパクトなスライディングウィンドウ型の断面潜在表現で形状を表す、位相を考慮した3D生成フレームワークSILSAを提案する。SILSAは高コストなボクセルトークンを生成する代わりに、3本の標準座標軸に沿った、重なり合う固定個数の断面を用いる。各トークンは奥行き方向の局所的な窓を要約し、断面間の連続性を保ちながら、単一段階のrectified flowによる生成を可能にする。 Slice VAEは向き付きの表面サンプルを複数軸の断面潜在表現に符号化し、疎な体積デコーダーで再構成する。一方、Volumetric Anchor Latticeは共通の3D作業空間を通して、方向別の断面ストリームを協調させる。構造の正しさを保つため、パーシステンス図を一致させ、隣接断面間のベッチ数の変化を整合させる、断面単位の位相的な教師信号を導入する。 実験では、SILSAは生成コストを大きく下げながら構造の忠実度を改善した。最も強力なベースラインに対し、PSNRを8.7%、カバレッジを絶対値で5.96ポイント、ベッチ誤差を9.2%改善する。また、次にコンパクトなベースラインよりトークン数を70.0%、疎なトークナイザーや階層型トークナイザーより98%超削減し、学習時メモリを40.4%、推論時間を58.5%削減した。定性的な結果でも、細い構造、繰り返し現れる構成要素、遠く離れた部分のつながりがよりよく保たれた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for thin or highly connected shapes. We introduce SILSA, a topology-aware 3D generation framework that represents shapes with compact sliding-window slice latents. Instead of generating expensive voxel tokens, SILSA uses a fixed set of overlapping slices along the three canonical axes, where each token summarizes a local depth window to preserve cross-sectional continuity and support single-stage rectified-flow generation. A Slice VAE encodes oriented surface samples into multi-axis slice latents and reconstructs them with a sparse volumetric decoder, while a Volumetric Anchor Lattice coordinates directional slice streams through a shared 3D workspace. To preserve structural correctness, we introduce slice-level topology supervision that matches persistence diagrams and aligns Betti transitions across neighboring slices. Experiments show that SILSA improves structural fidelity while substantially reducing generation cost. SILSA improves PSNR by $8.7\%$, coverage by $5.96$ absolute points, and Betti error by $9.2\%$ over the strongest baseline, while using $70.0\%$ fewer tokens than the next-most compact baseline and over $98\%$ fewer tokens than sparse or hierarchical tokenizers, effectively reducing training memory by $40.4\%$ and inference time by $58.5\%$. Qualitative results further show improved preservation of thin structures, repeated components, and long-range connectivity.
著者のコメント
Accepted at NeurIPS 2026. Project link: https://plan-lab.github.io/silsa
arXiv ID: 2610.02201 / 要約の誤りについて