A robust transcription system for soccer video database
Nhut Minh Pham, Duc Anh Duong, Quan Hai Vu · 2010
This paper presents a robust approach for the transcription of soccer video database. By exploiting audio channels in the video, spoken information is transcribed using a canonical speech recognition system. Since soccer videos vary in both speech quality and content, the transcription system is posed with three main problems: noisy data, foreign term interferences, and emotional variations in speech prosody. Three solutions are proposed to each of the problems respectively: a noise reduction scheme, a cross-lingual transliteration model, and an advanced acoustic modeling technique. Experimental evaluations of the proposed methods are conducted on the Vietnamese AFF Suzuki-cup database consisting of over 14-hour video. In the best case, system performance reaches 83.3% accuracy rate.