A Screen Content Video Quality Assessment Model Based on Adaptive 3-D Convolution

Hailiang Huang, Huanqiang Zeng, Jing Chen, Qi Lin, Huijie Zheng, Kai‐Kuang Ma · IEEE Transactions on Instrumentation and Measurement · 2025

This paper introduces the adaptive 3D convolution model (ATCM), a full-reference video quality assessment (VQA) model designed to evaluate quality degradation in distorted screen content videos (SCVs). SCVs are hybrid products of natural videos and computer-generated content, exhibiting diverse visual perceptions depending on regional complexity. For that, the 3D convolutional neural networks (3D-CNNs) are employed to convolve with the adaptive visual content, dynamically delving into the spatiotemporal intricacies of the SCVs. Specifically, the input SCVs are initially divided into complex and smooth sequences via the local video activity measure (LVAM), thereby constituting the adaptive visual content. Subsequently, the reference and distorted spatiotemporal features are extracted by the S3D networks, based on the Siamese networks. The quality scores of the distorted SCVs are then derived from the dual-channel spatiotemporal feature fusion. Notably, subjective evaluation labels are often absent in practical applications, the training relies on the initial pseudo-labels calculated from the classical full-reference algorithms, and gradually approximates subjective observations through iterative updates. Experimental results on two quality assessment databases for SCVs validate that the perceptual evaluations of the proposed ATCM align more closely with those made by the human visual system (HVS), compared to several classic and state-of-the-art algorithms.

Read the paper · More papers on PaperTik