Improving Multimodal Atmospheric Visibility Estimation Using Modern Image Feature Extractors and Attention-Based Fusion
Petr Doležel, Dominik Štursa, Dušan Kopecký, Jitka Kopecká · Zenodo (CERN European Organization for Nuclear Research) · 2026
Atmospheric visibility estimation is important for transport safety, environmental monitoring, and other operational applications. A previously proposed multimodal model combined RGB camera images with meteorological variables and showed that the fusion of visual and tabular data can improve visibility estimation. This study investigates whether the performance of this model can be further improved by modifying two components of the original architecture. First, six modern pre-trained image feature extractors are evaluated as replacements for the original EfficientNetV2M backbone. Second, three alternative multimodal regression heads are compared with the original fully connected fusion head. The experiments show that the choice of the image feature extractor strongly affects the global regression performance, while the fusion strategy further influences the practical reliability of the predictions. The best final configuration combines a ConvNeXt-based visual representation with a token-based attention fusion head. The results indicate that targeted adjustment of individual components of a multimodal visibility estimation model can lead to more favourable performance while preserving the same input modalities and evaluation protocol.