Audio-Visual Speech Recognition with Weighted KNN-based Classification in Mandarin Database

Tsang-Long Pao, Wenyuan Liao, Yu-Te Chen · 2007

Automatic speech recognition (ASR) by machine has been a goal and an attractive research area for past several decades. In recent years, there has been growing attractive research topic for overcoming certain audio-only recognition problems. Motivated by the multimodal nature of speech, the visual feature is considered to bring in information that dose not existing in the acoustic signal and enables improved system performance over audio-only methods. We first introduce the method for the extraction for the visual feature of the lip. In this paper, we compare three different weighting functions in weighted KNN-based classifiers to recognize ten digits, including 0 to 9, from Mandarin audio-visual speech. The classifiers studied include traditional KNN, weighted KNN, and weighted D-KNN. We also create a new audio-visual database in English and Mandarin. We will test this database for our proposed system, with some experimental results.

Read the paper · More papers on PaperTik