ReMove:Leveraging Motion Estimation for Computation Reuse in CNN-Based Video Processing

Masoumeh Khodarahmi, Mehdi Modarressi, Ardavan Elahi, Farhad Pakdaman · 2024

In this paper, we propose a method to reduce the computational load of Convolutional Neural Networks (CNNs) when processing video frames by exploiting computation reuse based on input similarity. Specifically, our approach leverages the temporal redundancy present in video sequences. In the existing computation reuse methods, if certain pixels of two consecutive frames are similar, the computations related to those pixels are skipped, and the results are reused from the previous frame. While pixel-wise comparison between consecutive frames can introduce overhead and partially offset computation reduction, we mitigate this by utilizing motion estimation information inherent in coded video frames. Motion estimation indicates whether a current block of the frame has already appeared in previous frames, allowing for direct reuse of computations without additional comparison overhead. Furthermore, we optimize by fusing CNN layers until the block size becomes smaller than the filter size, ensuring that not only the first layer’s computations but also multiple CNN layers’ computations are skipped. The experimental results demonstrate an average reduction of 35.4% in the computation amount of the VGG-16 CNN model with no significant loss in accuracy.

Read the paper · More papers on PaperTik