Fuzzy Inference based Pedestrian and Cyclist Tracking in a Compressed Video
S M Praveenkumar, Prakashgoud R. Patil, Sachin Karadgi, Prakash S. Hiremath · 2024
In recent times, the artificial intelligence (AI) based technologies are pervading the various human activities, for instance, smart manufacturing plants and autonomous driving vehicles (ADVs) on roads. With the advent of Industry 4.0, it is of common occurrence that human workers and autonomous mobile robots (AMRs) coexist and collaborate in the various tasks of smart manufacturing plants. The humans could be either walking or riding a bicycle in case of large area plants such as paper manufacturing industries. Thus, the industrial safety and traffic safety become cause of primary concern in the usage of AMRs and ADVs, respectively, which entails the need for equipping these intelligent devices with computer vision for detecting the humans, who are either walking or cycling, in their path of mobility, and thus avoid their collision. The integration of such detection systems is crucial for maintaining safe interactions between humans and machines. In the present study, a novel methodology based on fuzzy inferencing is designed for video analytics in the compressed domain, aimed at real-time detection and tracking of pedestrians and cyclists, by utilizing motion vectors (MV) and Transformer Prediction Heads-You Only Look Once version five (TPH-YOLOv5). The video compression formats MPEG-4 and H. 264 are considered. The proposed method can be adapted to any compression format and video camera settings without any prior setup. In the literature, many algorithms for tracking of pedestrians and cyclists in a compressed domain video are experimented, but these have some limitations such as (i) a fixed interval for keyframe occurrence, (ii) utilization of only $\mathbf{P}$ frames, and (iii) setup for only a single video codec. The proposed method is designed to address these limitations by employing keyframe intervals of variable length, utilizing P and B frames, and setting up two codecs MPEG-4 and H.264. Further the event box prediction in video frames is done by fuzzy inferencing. The benchmark video datasets, namely, Tsinghua-TDC, SCD, and KITTI datasets are considered for the experimentation. The ablation study has been carried out by choosing different variants of proposed method arising due to different video codecs, fuzzy representations, video datasets.