ELVSR: An efficient and lightweight convolutional neural network for FPGA-based FHD to UHD video super-resolution
Amirhossein Sadough, Mahyar Shahsavari, Mark Wijtvliet, Marcel A. J. van Gerven · Array · 2025
Video super-resolution (VSR) enhances mobile applications by overcoming mobile camera limitations, enabling data-saving where low-resolution images are upscaled locally to reduce bandwidth usage and improve performance in low-network environments, and improving gaming and video playback on high-resolution displays. However, existing VSR accelerators impose significant computational and memory burdens due to large neural networks, making real-time deployment on resource- and energy-constrained mobile devices impractical. To address this challenge, we propose an efficient and lightweight video super-resolution (ELVSR) model; a convolutional neural network designed for full-high-definition (FHD) to ultra-high-definition (UHD) video super-resolution. Implemented on field-programmable gate arrays (FPGA), our approach optimizes the trade-off between hardware efficiency and image quality, making it particularly suitable for mobile applications. Conventional FHD to UHD studies often evaluate quality using lower-resolution datasets, which can misalign with true efficacy and lead to unnecessary computational demands. To overcome this limitation, we utilize a 4K UHD dataset to assess quality performance directly at the target resolution. We also introduce an application-specific integrated circuit (ASIC)-oriented area-cost evaluation method for measuring the relative area cost of FPGA resources, which facilitates fair comparisons across designs and complements efficiency evaluations. Evaluations show that our work delivers effective image quality and optimal hardware efficiency, supporting [email protected] FPS with a power consumption as low as 2.557 watts, notably lower than state-of-the-art VSR accelerators. Our approach is particularly well-suited for mobile applications where resource and energy constraints are critical, improving energy-delay-area product by ∼ 2 × to ∼ 27 × over state-of-the-art accelerators.