Video quality metric based on fixation prediction and foveal imaging
Junyong You, Touradj Ebrahimi, Andrew Perkis · 2012
This paper proposes a full-reference video quality metric based on foveated vision mechanism. Due to a non-uniform distribution of photo-receptors on the retina, the human visual system (HVS) has the highest resolution around the fixation point of eyes and dramatically decreases away from this point. Two key factors in the foveated vision are fixation point and retinal eccentricities of different visual objects. Based on an advanced video attention model in quality assessment scenarios, eye fixations are predicted from the attention map using a winner-takes-all (WTA) neural network. Four quality features describing distortions on luminance, spatial and temporal activities, as well as chrominance are derived between foveated representations of reference and distorted video frames. These quality features are then combined together by an appropriate spatiotemporal pooling scheme to build a video quality metric. Experimental results with respect to publicly available video quality databases demonstrate that the proposed quality model outperforms a previously proposed foveated video quality metric as well as state-of-the-art video quality models.