Towards Efficient Neural Video Delivery via Scalable and Dynamic Networks
Shan Jiang, Haisheng Tan · 2024
With the advancement of deep learning-based super-resolution (SR), many video systems are opting to integrate SR models into the video delivery pipeline to achieve higher resolution, called neural video delivery. By transmitting SR models that are fine-tuned on the server for each video segment, the video quality has been significantly improved through executing SR models on the client. However, it introduces additional bandwidth resources for model transmission and computational costs for frame enhancement, which are bad for the overall performance of the system. To end this, we propose a novel video delivery framework via Space-Time Mixture of Experts (ST-MoE) to make the balance between them while maintaining video quality. Inspired by different requirements for the depth of SR models within intra-frame regions, we put forward Zero-padding Module to accelerate SR by routing different inputs to the appropriate sub-network. Besides, due to the varying similarity between inter-frame regions, we decompose our SR model into a shared portion for the whole video and the private portions for each video segment, aiming to decrease the parameters that need to be transmitted. Our experiments on diverse video streams show that our method achieves 40.98dB PSNR on average, which is 0.42dB better than the state-of-the-art. Besides, it reduces the transmission of SR models by 60% for each video segment and compresses the computation costs by 15% by distributing 45% of the frame patches to the tiny shared network.