MAE-CNN: A Multi-Scale Attention Enhanced Convolutional Neural Network for CU partition prediction

Junwei Chen, Chengji Zhao, Rongbang An · 2024

With the continual advancements in video encoding technology, the H.265 and H.266 standards have significantly reduced bit rates and enhanced data compression efficiency compared to their predecessors. H.265 employs a quadtree partitioning mechanism, while H.266 further develops this approach with the introduction of a Quad-tree with Nested Multi-type Tree (QTMT) partitioning system. These complex partitioning mechanisms require the exploration of all potential CU division options and the calculation of corresponding rate- distortion (RD) costs, substantially increasing encoding complexity. To address this challenge, this study introduces a Multi-scale Attention Enhanced Convolutional Neural Network (MAE-CNN), specifically designed to reduce the complexity of intra-frame video encoding. This deep convolutional neural network, trained using CU data collected through an official test model, incorporates the Inception module and the Convolutional Block Attention Module (CBAM). It significantly improves the detection of image details and broad area features, while also enhancing the network's ability to identify crucial feature information. Integrating this model into the encoder shows that under the full intra-frame encoding configuration, the algorithm saves an average of 64.07% of encoding time, with less average quality loss and lower bitrate growth, significantly better than traditional encoding methods. This research not only markedly improves encoding efficiency but also broadens the applications of deep learning in the video encoding field.

Read the paper · More papers on PaperTik