Mix-Based Training Strategies for Learning Implicit Neural Representations
Dongshen Han, Chaoning Zhang, Sheng Zheng, Fachrina Dewi Puspitasari, Yang Yang, Heng Tao Shen · IEEE Transactions on Multimedia · 2025
With coordinates as the input and RGB pixel values as the output, a neural network can be used to represent an image, which is widely known as Implicit neural representations (INRs). Previous works on INR have mainly focused on learning an invariant image target without exploring the impact of learning strategies on learning INR. It is observed that there is a substantial variation in PSNR among different images, and our preliminary investigation shows that, in the early training stage, learning complex image content yields significantly better performance than simple image content. Inspired by this finding, we conjecture that increasing INR task complexity in the early stage of training might boost INR performance and thus propose to intentionally contaminate the target image with another complex image. Our proposed method is called Mix-INR, which adopts a two-stage training to first learn a pseudo-target image (contaminated target) and then learn the real-target image (uncontaminated target). To generate the pseudo-target image, we experiment with two contamination methods (blending and replacement), both of which show superior performance and verify our conjecture. INRs have gained popularity as a promising approach for representing a variety of data types, including images of the task complexity of the pseudo-target image, we set the contamination image from a complex natural image to a random-noise image. Moreover, we propose a dynamic contamination method to smoothly transition from the pseudo-target image to the real-target image. Experimental results demonstrate that our proposed method achieves competitive performance, which suggests that INR can be improved by manipulating the task complexity in the early stage of training.