A Dual-Path Deep Learning Framework for Video Quality Assessment: Integrating Multi-Speed Processing and Correlation-Based Loss Functions
Hang Yu, Ruiming Tian, David Li · 2025
This paper presents a novel framework for video quality assessment (VQA) that builds upon the KVQ-Challenge platform, incorporating advanced deep learning techniques to improve accuracy in AI-driven video evaluation. Leveraging the SlowFast model architecture, our approach effectively captures fine-grained details and broader motion contexts by processing multi-speed visual streams. To enhance performance, we employ PLCC and Rank Loss functions to improve correlation accuracy and ranking precision, supported by adaptive learning rate scheduling and multiple correlation metrics for robust evaluation. The model architecture integrates components such as PatchEm-bed3D, WindowAttention3D, Semantic Transformation, Global Position Indexing, Cross Attention, and Patch Merging, each contributing to comprehensive feature extraction and aggregation. Extensive experiments on public datasets demonstrate that our model achieves superior results in both objective metrics and perceptual quality compared to existing methods, establishing a solid benchmark for future research in AI-based video quality assessment.