Shot Boundary Detection and Convolutional Neural Network for video Classification
Federico Candela, Angelo Giordano, Maurizio Campolo, Francesco Restuccia, Francesco Carlo Morabito · 2023
The classification of videos is a crucial computer vision task. Convolutional Neural Networks (CNN) have become a popular choice for video classification due to their ability to learn hierarchical representations of the input data. Shot boundary detection is another technique that is often used in video classification to segment the video into shots, which are temporal segments that contain a continuous action or scene. In the long-term video of a television broadcast there are continuous scene changes, some of which better define the beginning or end of a program, an advertisement, a type of human action and much more, supporting information for classification made only on specific video segments. Shot boundary detection can be incorporated into CNN models to segment the video into shots, and then feeding each shot into the CNN-based model for classification of television programs and their content. In this paper the implementation of a shot boundary technique will be described together with a pre-trained CNN on ImageNet for the video classification of some types of programs and television contents according to a schedule provided by the Italian national communications agency, the presented framework succeeds in achieving accuracy levels of 96%.