DP-CNN: Depth and Partition Convolutional Neural Network for VVC Intra Coding
Xu Li, Zongju Peng, Fen Chen, Hai Xiang, Qiong Liu · IEEE Transactions on Consumer Electronics · 2025
Multi-Type Tree (MTT) block partition technology introduced by next-generation Versatile Video Coding (VVC) significantly improves Rate Distortion (RD) performance. Nevertheless, the complicated recursive search for the optimal partition structure significantly increases the computational complexity of VVC. To address this issue, previous fast intra-coding methods were intended to expedite the coding process by predicting the partition types of a Coding Unit (CU). However, these methods are time-consuming and few methods directly predict the coding depth to reduce complexity. Therefore, in this paper, we propose a two-stage Depth and Partition Convolutional Neural Network (DP-CNN) for VVC. The method can efficiently predict the Quad Tree with a nested Multi-type Tree (QTMT) structure and coding depths, and effectively reduce unnecessary partition processes under the condition of lower distortion. First, a Depth Convolutional Neural Network (D-CNN) and a depth post-processing algorithm are designed to predict the coding depth of 64× 64 blocks, facilitating early termination of the RD optimization (RDO) process. We then use a Partition Convolutional Neural Network (P-CNN) and a partition decision algorithm to extract boundary features of 4× 4 blocks for the 32× 32 CU and its sub-CUs through a single inference process. The coding performance and complexity are balanced by the synergistic integration of the above methods. Meanwhile, the CNN inference process has been optimized to eliminate unnecessary delays to minimize additional time overhead. Comprehensive experiments are performed to validate the efficacy of the proposed DP-CNN. The experimental results show that the proposed method reduces the VVC coding time by approximately 38.86% 59.69% while maintaining a negligible Bjontegaard Delta Bit-Rate increase of 0.47% 1.50%, superior to the existing state-of-the-art methods.