Hardware Acceleration Model Construction and Implementation for Human Pose Recognition
Bei Xue · 2025
Human pose recognition is widely used in human-computer interaction, virtual reality and other fields, but the commonly used framework model is complex and the cost of GPU deployment is high. This paper uses lightweight PoseNet model combined with FPGA hardware devices to explore the method of hardware logic implementation of AI model. Deep separable convolution decomposes the traditional convolution into deep convolution and point-by-point convolution, which significantly reduces the computational complexity and parameters of PoseNet and improves the efficiency of hardware deployment. In the construction of the hardware model, the deep convolution and point-by-point convolution algorithms were first verified by Python. Then, the hardware model was developed in C++ in the HLS environment, and its performance was improved through data flow optimization, resource sharing, and array partition memory strategies. The experimental results show that the optimization of point-by-point convolution and deep convolution can greatly improve the performance of parallelism, delay and loop times. The parallelism of point-by-point convolution is increased by 384 times, the parallelism of deep convolution is increased by 9 times, and the IP core resource consumption of the hardware accelerator is reasonable, which realizes efficient real-time human posture recognition.