Diffuse and Refine Latent Prior with Transformers For Neural ISP

Zhipeng Mo, Wenbo Li, Sihao Ding · 2025

The goal of neural ISP is to convert the RAW data captured by camera sensors into display-referred sRGB images via an end-to-end neural network. Most methods are based on the regression, so they are prone to recovering images with fewer details due to the constraint of the regression loss. Some methods attempt to recover more details by resorting to GAN or diffusion model (DM) at the cost of hurting the signal-to-noise ratio and image structures. To circumvent such a trade-off, inspired by the recent progress in image deblurring, we propose a baseline that integrates the regression-based Transformer and DM. Specifically, DM is performed in a highly compact latent space to predict the latent prior that mimics the ground-truth latent prior compressed from the target sRGB image; then, the predicted latent prior is injected into the Transformer to guide the RAW-to-sRGB mapping process. Despite the superior performance achieved by this baseline, we conjecture that the latent prior predicted by DM is not good enough to mimic the ground-truth latent prior. Therefore, based on the baseline, we propose D&R in which the latent prior is diffused and refined with Transformers. The experimental results show that our D&R achieves state-of-the-art performance across multiple metrics on ZRR dataset and SID dataset, and also demonstrates the superiority over the baseline.

Read the paper · More papers on PaperTik