A hybrid approach for Discourse Segment Detection in the automatic subtitle generation of computer science lecture videos
Rajeswari Sridhar, S Aravind, Hamid Muneerulhudhakalvathi, M Sibi Senthur · 2014
The aim of this paper is to develop an automatic subtitle generation system for computer science lecture videos. CMU Sphinx Speech API is used to accomplish speech recognition. The main challenge of this work, is to align the translated text with the video. Discourse Segment Detection (DSD) is the process of analyzing and identifying discourse boundaries in human speech. Discourse Segment Detection (DSD) is carried out that classifies word boundaries and groups words until a discourse break occurs. The approach that has been devised in this paper for DSD to identify word boundary is a hybrid approach combining acoustic and linguistic features from the speech. This helps to segment the text obtained from Speech Engine, group words that can be written to the subtitles file without violating the subtitle standards. The devised approach has shown an improved performance than the existing approach as the error has reduced from 30% to 18 %.