Incorporating Frame Image and Frame Sequence into Ensemble Learning Networks to Improve the Accuracy of Physical Bullying-Detecting Model
Runqian Wang · IOP Conference Series Materials Science and Engineering · 2019
Abstract In this research, a convenient, reliable, and fast system is built for violence detection in schools. This system includes two independent model designed separately for detecting violence in videos and pictures. This ensemble learning structure allows the system to be more accurate and less dependent on the background, or, minimize the influence of changing circumstances. The reason is that the image-identifying network did better on classifying the image, while the video-identifying network can minimize the influence of the background by using optical flow. Therefore, by combining them together the total accuracy increases and the dependency on environment is minimized. The video-identifying model was based on Darknet19, ResNet, optical flow and LSTM. Optical flow could eliminate the background’s influence by extracting moving features. The image-identifying model was based on ResNet. ResNet was tested and selected from 3 different popular networks due to its high accuracy. The system was trained and tested on data from various datasets online as well as different videos and pictures from multiple sources. The final system achieves an accuracy of 90.90% from multiple-sourced videos, and is visualized for a more straight-forward result. The final test shows that the system not only has a high accuracy, but also can process data with an astonishing speed of 75 FPS (frames per second), which is about 3 times faster than normal videos. All of these imply a high practical value of the final system.