Long Short-term Convolutional Transformer for No-Reference Video Quality Assessment
Junyong You · 2021
No-reference video quality assessment has not been widely benefited from deep learning, mainly due to the complexity, diversity and particularity of modelling spatial and temporal characteristics in quality assessment scenario. Image quality assessment (IQA) performed on video frames plays a key role in NR-VQA. A perceptual hierarchical network (PHIQNet) with an integrated attention module is first proposed that can appropriately simulate the visual mechanisms of contrast sensitivity and selective attention in IQA. Subsequently, perceptual quality features of video frames derived from PHIQNet are fed into a long short-term convolutional Transformer (LSCT) architecture to predict the perceived video quality. LSCT consists of CNN formulating quality features in video frames within short-term units that are then fed into Transformer to capture the long-range dependence and attention allocation over temporal units. Such architecture is in line with the intrinsic properties of VQA. Experimental results on publicly available video quality databases have demonstrated that the LSCT architecture based on PHIQNet significantly outperforms state-of-the-art video quality models.