Multi-modal Vit-Wavenet: A novel algorithm for precise identification and classification of image noise
Maddimsetty Surya Prakash, Tamilselvan Sadasivam · 2025
Accurate identification and characterization of image noise are imperative for ensuring the reliability and effectiveness of image processing tasks. In this context, the development of deep learning models capable of classifying various types of image noise holds significant importance in understanding and mitigating its adverse effects on image quality and analysis. In this study, we introduce Multi-Model ViT-WaveNet, to develop a deep learning model capable of classifying and characterizing various types of image noise, including Gaussian, Poisson, salt-and-pepper, and more. By using the power of Vision Transformers (ViTs) and wavelet transforms, ViT-WaveNet offers a comprehensive approach to understanding the nature and extent of noise present in an image. By integrating spatial and frequency-domain features extracted by ViTs and wavelet transforms, respectively, our algorithm achieves enhanced performance in noise identification tasks. Through multi-modal feature fusion, ViT-WaveNet captures intricate patterns and relationships within image noise, providing valuable insights into its characteristics. The proposed algorithm contributes to advancing image processing techniques, with potential applications in medical imaging, satellite imagery analysis, digital photography, and more. Experimental results demonstrate the effectiveness and versatility of ViT-WaveNet in addressing the challenges of image noise identification and characterization, paving the way for improved image quality and decision-making processes in various domains.