Ultra-Low Complexity Neural Networks for Next Generation Video Decoding

Kiran Misra, Shashwat Ranjan Chaurasia, Andrew Segall, Byeong-Doo Choi · 2025

We consider the problem of embedding a neural network directly into a video decoder. This requires a design with complexity suitable for implementation on mobile and power constrained devices. To achieve this goal, we explored Multi-scale CNN (MSCNN) design in [1]. In this paper, we improve the design to support super resolution spatial scale factors SF==(1.5×, 2×, 3×, 4×, 6×) by modifying the polyphase filter (Figure 1a) that generates an upsampled output using g(scale) phases and stride of Sscale. When SF= 1.5 ×, g(scale, Sscale) = (9,2); Otherwise it is (scale2,1). gsG, kK, and sS denote channel group size of G, kernel size of K×K, and stride of S. To reduce per-pixel Multiply-Accumulates (MACs), the 3×1 and 1×3 convolutional layers use Canonical Polyadic (CP) decomposition and reduced channel count. These changes reduce MACs/pixel from 1,924 in [1] to 1,192 to 584. Figure 1b, shows the placement of MSCNN in AVM [2]. We code 4K video, using AOMedia's Adaptive Streaming (AS) test conditions and compare MSCNN versus following resampler combinations: Downsampling - [L5: Lanczos(5), L6: Lanczos(6)]; Upsampling - [L5, L6, BL: Bilinear, BC: Bicubic]. We observe MSCNN provides on average 30.4% rate saving.

Read the paper · More papers on PaperTik