Efficient Implicit Neural Representations for Videos with Feature Modulation

Haeyoon Yang, Nam Ik Cho · 2025

Implicit neural representations for videos (NeRV) have gained attraction due to their ability to embed videos into neural networks effectively. A major challenge in implicit neural representations for videos is achieving high performance with small-sized networks. Therefore, it is crucial to minimize the intrinsic redundancy of the videos embedded within the network. In this paper, we propose a new method called ModNeRV (NeRV with feature modulation), which aims to reduce redundancy using base grids and feature modulation. The base grids, structured with multiple spatial resolutions, capture the general content of the video that is shared across frames. Frame-Specific features are then generated through frame-wise feature modulation applied to these base grids. Additionally, feature modulation is further utilized within ModNeRV blocks to generate details specific to each frame. As a result, ModNeRV demonstrates superior performance compared to previous NeRV-based methods in video reconstruction. Specifically, on the "Big Buck Bunny" and UVG datasets, ModNeRV achieves an increase in PSNR of 9.5 dB and 2.5 dB over NeRV with model sizes of 0.75M and 3M, respectively. Furthermore, in terms of video compression, ModNeRV outperforms both previous NeRV-based methods and traditional compression techniques such as H.264 (AVC) and H.265 (HEVC).

Read the paper · More papers on PaperTik