Construction vision improved with paired synthetic dataset under bad conditions

A paired synthetic construction-site image dataset for robust computer vision under adverse conditions

Computer Vision and Pattern Recognition

Summary

Monitoring construction sites with computer vision can be tricky when weather or lighting is poor, but most existing image collections don’t include these tough conditions. The authors created ConSynth-X, a large set of synthetic construction images matched to real ones, showing effects like rain, fog, and night scenes. This matching lets people test how well computer vision systems handle different problems in a controlled way. The dataset also includes helpful labels for tasks like detecting objects or answering questions about the images.

What this means in practice

Authors

Viet Huy Duong, Ruoxin Xiong, Md Abdullah Al Forhad, Weishi Shi

Abstract

Computer-vision systems used for construction monitoring can degrade under adverse environmental and visual conditions, yet such conditions remain underrepresented in existing construction image datasets. We present ConSynth-X, a paired synthetic construction-site image dataset containing 34,199 images derived from 3,109 real-world source scenes. The dataset comprises 11 condition-specific subsets spanning precipitation, fog, nighttime illumination, adverse weather at night, and small-object or long-distance views. Each synthetic image is linked to its corresponding source scene, enabling controlled comparison across environmental and visual conditions. ConSynth-X includes source-derived annotations, generation metadata, provenance information, and image-quality indicators, supporting object detection, image captioning, visual grounding, and visual question answering. Technical validation evaluates source-synthetic fidelity and alignment with real adverse-condition imagery using embedding-based similarity and distributional analyses. The dataset provides a structured resource for evaluating and improving the robustness of construction vision and vision-language models under challenging field conditions.