Dataset Pipeline Evaluation Gallery Human Eval Download

A Construction Dataset Built for Field Conditions

Examples of ConSynth-X clear, fog, rain, snow, and small-object condition variants.
Condition variants generated from construction scenes.
34,199 released images
11 synthetic conditions
3 source datasets

Released field schema shared across the dataset

Field Type Description
image binary JPEG byte stream stored in Apache Arrow format; pass-through with no re-encoding.
image_id string Unique image identifier within a sub-dataset.
source_id string | null Identifier of the original clean image, enabling linkage to source records.
source_dataset string Source dataset identifier: cs10k, soda_voc, or soda_ktsh.
condition string Primary condition label.
condition_labels list[string] Set of applicable condition labels for multi-label filtering or stratified analysis.
objects list[struct] Object annotations containing class_id, class_name, and normalized bbox coordinates.
pipeline struct Synthesis metadata: method, checkpoint, prompt template, and parameter record.
quality_scores struct | null Optional metrics: dino_sim, ssim, clip_sim, and lpips.
quality_alert bool | null Quality flag where dino_sim < 0.75; null if unavailable.

Number of records by synthetic condition

Horizontal bar chart of ConSynth-X record counts by synthetic condition.

11 Conditions
Across Four Categories

11 condition labels
4 taxonomy groups
7 weather variants
7 labels

Weather

rain_light rain_heavy snow_light snow_heavy fog_light fog_medium fog_heavy
1 label

Lighting

night
2 labels

Compound (Night × Weather)

night_rain night_snow
1 label

Scale / Distance

small

Built on Open Construction Datasets

Construction Site 10K

Chen, X. and Zou, Z. (2026). Are large pre-trained vision language models effective construction safety inspectors? Data-Centric Engineering, 7, e11.

BibTeX
@article{chen2026large,
  title={Are large pre-trained vision language models effective construction safety inspectors},
  author={Chen, Xuezheng and Zou, Zhengbo},
  journal={Data-Centric Engineering},
  volume={7},
  pages={e11},
  year={2026},
  publisher={Cambridge University Press}
}
SODA Dataset

Duan, R., Deng, H., Tian, M., Deng, Y., and Lin, J. (2022). SODA: A large-scale open site object detection dataset for deep learning in construction. Automation in Construction, 142, 104499.

BibTeX
@article{duan2022soda,
  title={SODA: A large-scale open site object detection dataset for deep learning in construction},
  author={Duan, Rui and Deng, Hui and Tian, Mao and Deng, Yichuan and Lin, Jiarui},
  journal={Automation in Construction},
  volume={142},
  pages={104499},
  year={2022},
              publisher={Elsevier}
}
SODA-ktsh Dataset

Deng, H., Fu, K., Yu, B., Li, H., Duan, R., Deng, Y., and Lin, J. (2025). Enabling high-level worker-centric semantic understanding of onsite images using visual language models with attention mechanism and beam search strategy. Buildings, 15(6), 959.

BibTeX
@article{deng2025enabling,
  title={Enabling High-Level Worker-Centric Semantic Understanding of Onsite Images Using Visual Language Models with Attention Mechanism and Beam Search Strategy},
  author={Deng, Hui and Fu, Kejie and Yu, Binglin and Li, Huimin and Duan, Rui and Deng, Yichuan and Lin, Jia-rui},
  journal={Buildings},
  volume={15},
  number={6},
  pages={959},
  year={2025},
  publisher={MDPI}
}