DSNet-MV: Fast summarization of Surveillance Video’s using Deep learning in Compressed Domain using Motion Vectors
Lakshya Lakshya, Shivaank Agarwal, Venkata Suneel Kota, Mallikarjuna Rao Voleti · 2021 IEEE 18th India Council International Conference (INDICON) · 2021
Video summarization is an attractive choice to handle the large bulk of data that is produced by a home surveillance system. However, the video summarization algorithm should be faster than the video stream generation rate for effectiveness. Moreover, there are privacy concerns when a potentially sensitive video stream is processed. To alleviate these issues, Motion Vector(MV) information in the encoded stream can be used instead of decoded RGBffUV pixel values of the frames. However, most of the recent Deep Learning(DL)-based technologies for video summarization use pixel-level information. In this paper, we try to bridge this gap by modifying pixel domain-based DL techniques to use MV data. Specifically, we alter DSNet by introducing a Masked Multi-Headed Attention layer and an aggregation module. Moreover, we explain how training DSNet using Masked Multi-Headed Attention allows us to process live streams of data, as compared to stored clips. Finally, we show that even just by using MV information, impressive performance is achieved by our modified DSNet model on TVSum and SumMe dataset.