Implementation of parallel full search algorithm for motion estimation on multi-core processors
Jing Zhou, Liangbao Jiao, Xuehong Cao · 2011
In order to speed up H.264/AVC coding efficiency, this paper proposed a parallelization approach of full search (FS) algorithm for motion estimation on Graphic Processor Unit (GPU) using computing unified device architecture (CUDA). According to the independence among different macro-blocks (MBs), we mapped the traditional sequential FS algorithm for motion estimation to CUDA parallel computing model with optimizing memory usage, taking full advantage of the powerful parallel computing capability to speed up FS motion estimation. Experimental results show that our implementation on CUDA demonstrates substantial improvement up to 50 times than CPU counterpart available and can effectively speed up the FS for motion estimation.