arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

建築図面でハエ視覚回路と学習済みモデルを比較

What Survives on Real Drawings: Active Sampling, Connectome Wiring, and Matched Baselines in Architectural Document Vision

Dmitry Kuklev

この論文をやさしく読む

ひとことで言うと

建築図面の認識で、ハエの視覚神経回路を模した小型モデルと、情報量をそろえた学習済みモデルを比べた。

何に役立つ?

図面認識でノイズへの頑健性、モデル規模、学習の有無を比較する材料になる。ハエ型回路の実用上の優位が全条件で示されたわけではない。

この研究の面白いところ

きれいな合成データではCNNが高精度だが、14枚の実務図面では小型ハエモデルの平均適合率0.505が大きなネットワークの0.415を上回った。一方、元の回路は動く網膜のみの構成には勝てなかった。

どこまで分かった?

実務図面での評価は14枚であり、事前登録した保留分割では学習済みCNNが最良だった。条件によって順位が変わる。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ハエの視覚系の神経接続に制約を課し、動きに対して最適化した後に固定したモデルは、定められた動きで建築図面を読み取ると、テクスチャの表現として使える。本研究は、同じ721個の光受容器標本を見るよう情報量を合わせた基準モデルと比較する。きれいな合成データでは固定モデルからの転移は可能だが、課題に合わせた学習には及ばない。1例からのハッチング照合で面積加重精度は0.857に対し、5,888パラメータのCNNは0.959だった。壁の領域分割でのIoUは0.619に対し、対応するネットワークは0.905だった。 スキャンノイズと太くなった線の下では、学習したネットワークの精度が最大0.188低下した一方、固定した処理系列の低下は0.030にとどまった。一度だけ開いた14枚の実際の業務用図面では、1,876パラメータのハエモデルの平均適合率は0.505で、約200倍大きなネットワークの0.415を上回った。事前登録された保留データ分割では、きれいなデータでの順位が再確認された。回路は0.835、光受容器のみは0.894、学習済みCNNは0.971~0.980だった。 神経接続の次数または型の組と神経伝達物質の符号を保ちながら接続を組み替えると、3つの乱数シードで精度が0.271~0.356低下し、接続の正確な構造が重要だと分かる。ただし、元の回路は動く網膜だけの構成には勝てず、T4/T5を停止しても両課題の性能は保たれた。観測を長くすると、刺激が一周期に及んだ時点で回路と光受容器の順位が逆転するが、これはT4/T5によるものではない。したがって、能動的な標本取得と接続構造は重要であるものの、きれいなデータでの実用性能は課題に合わせて学習したネットワークが主導し、有用な転移上の利点の多くは網膜部分に由来する。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-22 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

A connectome-constrained model of the fly visual system, optimized for motion and then frozen, can be driven over architectural drawings by prescribed motion and used as a texture representation. We compare it with information-matched baselines that see the same 721 photoreceptor samples. On clean synthetic data the frozen model transfers but loses to task training: 0.857 area-weighted accuracy in one-shot hatch matching versus 0.959 for a 5,888-parameter CNN, and 0.619 IoU in wall segmentation versus 0.905 for a matched network. Under scan noise and thickened strokes, the trained networks lose up to 0.188 accuracy while the frozen pipeline loses 0.030. On fourteen production sheets, opened once, a 1,876-parameter fly model reaches 0.505 average precision versus 0.415 for a network two hundred times larger. A preregistered held-out split confirms the clean-data ordering: 0.835 for the circuit, 0.894 for receptors only, and 0.971-0.980 for trained CNNs. Rewiring the connectome while preserving degrees or type pairs and transmitter signs costs 0.271-0.356 accuracy across three seeds, so the exact wiring is load-bearing. Yet the intact circuit does not beat its moving retina, and T4/T5 silencing leaves both tasks intact. Longer observations reverse the circuit-receptor ordering once the stimulus spans a period, but not through T4/T5. Thus active sampling and exact structure matter, while clean-data practical performance remains dominated by task-trained networks and the useful transfer margin is largely retinal.

著者のコメント

11 pages, 8 figures, 6 tables. Major revision with a preregistered held-out evaluation, matched graph nulls, temporal sampling controls, real-drawing evaluation, and an ancillary animation

arXiv ID: 2609.24565 / 要約の誤りについて