Spatiotemporally Modulated Dual-MLP Architecture for Neural Video Representation and Compression

Guoxin Wu, Cih-Wei Wong, Hsu-Feng Hsiao · 2025

This paper presents a novel modulated dual-MLP framework for efficient neural video representation and compression using implicit neural representations. Our approach overcomes limitations in existing methods by integrating a modulator network with sinusoidal activation functions within a dual-MLP architecture. Compression performance is further improved through advanced techniques such as layer-adaptive pruning, quantization, and context-adaptive binary arithmetic coding. Pruning is applied independently to each layer, and our pipeline retains all necessary metadata, enabling accurate recovery of compressed weights and model reconstruction for thorough evaluation. Additionally, our framework supports color space conversion; fine-tuning RGB-trained models with YUV420 data allows seamless transitions with minimal architectural changes, which is beneficial for digital broadcasting and streaming, where YUV formats are common. Extensive experiments on the challenging UVG dataset show that our method achieves up to 0.89 dB higher PSNR than leading INR approaches at 12.5 million parameters, while also delivering improved rate-distortion performance compared to other INR methods and popular codecs like AVC and HEVC.

Read the paper · More papers on PaperTik