欠陥の視覚的証拠を保つ産業用異常検出
Industrial Anomaly Detection via Defect-Grounded Reasoning in Visual Latent Space
この論文をやさしく読む
ひとことで言うと
欠陥の位置に関する画像上の証拠を保ちながら、産業画像の異常を説明する方法。
何に役立つ?
製造物の画像点検で、欠陥を検出するだけでなく根拠を追って説明する仕組みの検討に役立つ。
この研究の面白いところ
視覚的な潜在空間で証拠を段階的に精密化し、2万件超の画像・質問データも作った。
どこまで分かった?
性能は要旨にいう同程度の規模の比較手法に対する結果で、具体的な指標値は記載されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
産業用の異常検出は、従来の検出と位置特定から、細かな欠陥を記述、説明、推論できるマルチモーダル点検へ発展している。最近のマルチモーダル大規模言語モデルを使う方法は文章による推論と視覚的な手掛かりで理解を深めるが、細かな点検では2つの課題がある。局所的な画像領域を繰り返し見直したり追加の道具を使ったりする必要があり、得られた欠陥の局所的な証拠が後の推論で確実に保持されない場合がある。著者らはAnomaly-LRという、欠陥に根差した潜在表現での推論方法を提案する。まず入力の全体像を捉え、その後、異常に関係する表現を視覚的な潜在空間で直接、段階的に精密化する。また、潜在表現での推論向けに設計した初の産業用異常検出指示データセットIAD-LR-22Kを作った。4,523枚の産業画像から、22,228件の画像と質問の組を収め、全体の文章による推論履歴と領域単位の視覚的注釈を備える。複数のベンチマークでの広範な実験では、外部の参照や道具を必要とせず、同程度の規模の方法の中で最新水準の性能を得た。コードとデータは公開予定である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Industrial anomaly detection (IAD) is evolving beyond conventional detection and localization toward multimodal inspection systems that can describe, explain, and reason about fine-grained defects. Although recent multimodal large language model (MLLM)-based methods improve anomaly understanding through textual reasoning and visual guidance, they face two limitations in fine-grained inspection. First, their visual refinement often requires iteratively revisiting local image regions or augmenting with additional tools. Second, the resulting local defect evidence may not be reliably preserved throughout subsequent reasoning. To address these, we propose Anomaly-LR, a defect-grounded latent reasoning framework that first forms a global understanding of the input and then progressively refines anomaly-relevant representations directly in the visual latent space. We further construct IAD-LR-22K, the first IAD instruction dataset designed for latent reasoning, containing 22,228 image-question instances from 4,523 industrial images, with global textual reasoning traces and region-level visual annotations. Extensive experiments show that Anomaly-LR achieves state-of-the-art performance among comparable-scale methods across multiple IAD benchmarks, without requiring external references or tools. The code and data will be released at https://github.com/Yen666/Anomaly-LR.
著者のコメント
5 pages
arXiv ID: 2609.29457 / 要約の誤りについて