Optical image processing and applications empowered by vision-language models

J. Xiao, Zhe Sun, Hongjun An, Haofei Zhao, Maosheng Qiu, Xuelong Li · iOptics · 2025

Optical images present significant analytical challenges due to their high-dimensional structures, complex modalities, and multi-scale characteristics. This review systematically examines the technological evolution in optical image processing and applications. It traces the developmental trajectory from traditional image processing methods to deep learning approaches, culminating in the emergence of Vision Language Models. The discussion is organized around four pillars, including high-dimensional data representation, cross-modal feature fusion, semantic alignment, and reasoning strategy design. It synthesizes recent advances while mapping persisting constraints in network architecture, data dependence, and inference efficiency. The review closes by outlining priority research avenues aimed at elevating efficiency, generalization, and deployability in optical image understanding.

Read the paper · More papers on PaperTik