arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

悪天候や夜間の建設現場を比較できる合成画像集

A paired synthetic construction-site image dataset for robust computer vision under adverse conditions

Viet Huy Duong, Ruoxin Xiong, Md Abdullah Al Forhad, Weishi Shi

この論文をやさしく読む

ひとことで言うと

同じ建設現場を、霧、夜、悪天候などに変えた合成画像をそろえ、条件の違いで画像認識がどう変わるか比較できるようにしたデータ集です。

何に役立つ?

建設現場向けの物体検出や画像質問応答を、難しい見え方で評価する用途があります。元場面との対応があるため、場面自体の違いを抑えた比較ができます。

この研究の面白いところ

34,199枚を独立な場面として集めたのではなく、3,109の実場面から条件違いの画像を作っています。生成情報や来歴、画質指標も持たせています。

どこまで分かった?

悪条件の画像は合成であり、元の実場面数と画像総数は異なります。実画像との整合性の検証はありますが、現場での認識率改善の具体値は要旨に記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

建設現場の監視に使うコンピュータービジョンシステムは、悪い環境条件や視覚条件の下で性能が低下し得るが、既存の建設画像データセットにはこうした条件が十分含まれていない。本研究では、実際の3,109場面から作成した34,199枚の画像を含む、元画像との対応付き合成建設現場画像データセットConSynth-Xを提示する。 データセットは、降水、霧、夜間照明、夜間の悪天候、小さな物体または遠距離の視点を含む、条件別の11の部分集合からなる。各合成画像は対応する元の場面と結び付けられており、環境・視覚条件を変えた制御された比較ができる。ConSynth-Xには、元画像由来のアノテーション、生成メタデータ、来歴情報、画質指標が含まれ、物体検出、画像説明文の生成、視覚的グラウンディング、視覚的質問応答を支援する。 技術的検証では、埋め込みに基づく類似度と分布の分析を用いて、元画像と合成画像の忠実性、および実際の悪条件下の画像との整合性を評価する。このデータセットは、厳しい現場条件での建設分野の視覚モデルや視覚言語モデルの頑健性を評価・改善するための、構造化された資料を提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Computer-vision systems used for construction monitoring can degrade under adverse environmental and visual conditions, yet such conditions remain underrepresented in existing construction image datasets. We present ConSynth-X, a paired synthetic construction-site image dataset containing 34,199 images derived from 3,109 real-world source scenes. The dataset comprises 11 condition-specific subsets spanning precipitation, fog, nighttime illumination, adverse weather at night, and small-object or long-distance views. Each synthetic image is linked to its corresponding source scene, enabling controlled comparison across environmental and visual conditions. ConSynth-X includes source-derived annotations, generation metadata, provenance information, and image-quality indicators, supporting object detection, image captioning, visual grounding, and visual question answering. Technical validation evaluates source-synthetic fidelity and alignment with real adverse-condition imagery using embedding-based similarity and distributional analyses. The dataset provides a structured resource for evaluating and improving the robustness of construction vision and vision-language models under challenging field conditions.

著者のコメント

21 pages, 7 figures, 7 tables. Dataset and code are publicly available

arXiv ID: 2609.24075 / 要約の誤りについて