Initial Experiments with Tamil LVCSR
Melvin Jose J., Ngoc Thang Vu, Tanja Schultz · 2012
In this paper we present our recent efforts towards building a large vocabulary continuous speech recognizer for Tamil. We describe the text and speech corpus collected to realize this task. The data was complemented by a large amount of text data crawled from various Tamil news websites. The Tamil speech recognition system was bootstrapped using the Rapid Language Adaptation scheme which employs a multilingual phone inventory. After initialization, we built a word-based and syllable-based system with a Syllable Error Rate (SyllER) of 29.30% and 34.16%, respectively. We propose a data-driven approach to obtain better dictionary units to overcome the challenge of the agglutinative nature of Tamil. The approach produced a significant improvement of 27.20% and 15.12% relative SyllER on the test set over the syllable- and word-based systems, respectively. Our current best system has a SyllER of 17.44% on read newspaper speech.