2.8: A 210fps Image Signal Processor for 4K Ultra HD True Video Super Resolution

Ying‐Sheng Lin, Jun Ivan Nishimura, Chia‐Hsiang Yang · 2025

Video super-resolution (VSR) aims to convert low-resolution (LR) videos to high-resolution (HR) videos with high image quality [1]. It can be used for various video applications, such as streaming and conferencing, given low-bandwidth connectivity. Figure 2.8.1 shows a neural network (NN)-based VSR framework that comprises alignment and refinement stages [2]. Optical flow, which indicates the motion of the pixels, is used to align the objects between two frames and to generate a coarse HR frame. The coarse HR frame is then refined to produce the final HR frame. Videos are usually encoded in a compact frame-based format for efficient storage and transmission. Frames are categorized into several types (I, P, and B): I frame is encoded without referencing other frames, P frame is encoded by referencing previous frames, and B frame is encoded by referencing both previous and future frames. In each frame, blocks are also encoded by block prediction to reference similar blocks in other frames. The location of the reference block is indicated by a motion vector and the pixel difference is encoded as a residual. Dedicated processors have been proposed to deliver high throughput for VSR [3]–[5]. However, prior works treat video as a collection of independent images [3]–[4] or raw video [5], without considering the inter-dependency of frames. In the multi-image workflow shown in [5], extracted features of adjacent frames are reused, but temporal and spatial redundancies embedded in video frames are neglected for VSR acceleration. This work demonstrates an image signal processor for true video super-resolution with high frame rate and energy efficiency.

Read the paper · More papers on PaperTik