Automatic Video summarization with Timestamps using natural language processing text fusion

Ahmed Emad, Fady Bassel, Mark Refaat, Mohamed A. Abdelhamed, Nada Shorim, Ashraf AbdelRaouf · 2021

In videos, description and keywords play an important role in the choosing process of the right video to watch. The main idea of the proposed approach is to generate descriptions and timestamps for videos automatically. Our approach plays an essential role in reducing the time consumed searching for the proper video. It aims to save time for users watching wrong unwanted videos and saves their time using timestamps. Timestamps would help to find and watch only the desired part of the video. One of the main goals of our approach is actual keyword extraction. Extracted keywords help finding videos with the significant video's keywords. The summarizing of the video depends on frames, emotions and speech. Firstly the video content appears in the frame and output a summarized text for the video content. Secondly, emotion and how it changes during a specific period merged with the outputted summarization of the frames. Thirdly, the audio transcribing into text occurs and output an abstractive summarization of the audio track. Finally, the fusion happens between all summarizations (audio, video, emotion) using natural language processing techniques. Techniques such as tokenization, sentence segmentation and lemmatization & stemming, and then abstractive summarization. Video summarization occurs to get a meaningful accurate description of the video. Having an accurate description helps finding the inquired content matching the description. The implemented experiment showed that on average 87% of the participants found generated text well representing the video.

Read the paper · More papers on PaperTik