Further results on the recognition of a continuously read natural corpus

L.R. Bahl, R. Bakis, Paul S. Cohen, A. Cole, F. Jelinek, Blaine Lewis, R. L. Mercer · 2005

Further results have been obtained on the recognition of continuously read sentences from a natural language corpus of laser patents. The vocabulary is limited to the 1000 most frequently occurring words in the corpus. Our model of the task language has a perplexity of 24.1 words (corresponding to an entropy of 4.6 bits/word). This paper describes modifications and improvements to the system which have resulted in the lowering of the word error rate from the previously reported 33.1% to 8.9%.

Read the paper · More papers on PaperTik