Applied YOLOv5 and Context Constraint for Real-time and High Accuracy Human Detection

Van-Hung Le, Thi-Loan Pham, Hai-Yen Tran, Tien‐Thanh Nguyen · 2022

Detecting human in images or videos has been of interest to the computer vision research community for many years. This problem is also widely applied in areas such as security, tracking people, estimating posture, etc. In particular, there have been many impressive results in detecting people using CNNs as Fast R-CNN, VGG, SSD, Mask R-CNN, Mobilenet, etc. However, these studies are often applied on computers with high GPU speed, which makes building systems to detect, identify and track people expensive. The YOLOv5 was born with the purpose of running on low-cost, low-profile computers and training online from user data. In this paper, we propose the combination of YOLOv5 and context constraints for fast and accurate detection of people in images. The result of this study is a prepossessing step of 2D and 3D human pose estimation, evaluated and compared with other methods on the Human 3.6M dataset. For the evaluation we also mark the person with a bounding box on the test data of the Human 3.6M dataset (Protocol #1 - Subject 9, Subject 11). The person detection accuracy is$AP_{50}=99.78\%,AP_{55}=99.38\%,AP_{60}=98.42\%,AP_{65}= 97.07\%,AP_{70}=94.16\%$, respectively and the processing time is 55 Hz on the GTX 970 GPU 4G. Experimental results is available.

Read the paper · More papers on PaperTik