Video retrieval using speech and text in video
N. Radha · 2016
This paper presents an analysis process of lecture video retrieval with automated video indexing and video search for lecture video databases. The video retrieval extracts the relevant metadata from the two main parts of lecture videos, namely the visual screen and audio tracks. From the visual screen, we firstly detect the slide transition and extract each unique slide frame with its temporal transition considered as the video segment. The textual metadata from slide frames is extracted and then recognized using video OCR technique. Based on OCR results, the corresponding text in video information is saved. Secondly, the speech-to-text analysis is carried out from the audio signal. Sphinx speech recognition models are used for recognition of speech-to-text conversion process. In the proposed work, the combined text information from the visual slide frame and audio signal were used to retrieve the video from the lecture video database. The performance of this combined video retrieval system shows a significant improvement in performance when compared with an individual system built using audio and text in video systems respectively.