Improving Pronunciation Inference using N-Best List, Acoustics and Orthography
Gopala K. Anumanchipalli, Mosur Ravishankar, Raj Reddy · 2007
In this paper, we tackle the problem of pronunciation inference and out-of-vocabulary (OOV) enrollment in automatic speech recognition (ASR) applications. We combine linguistic and acoustic information of the OOV word using its spelling and a single instance of its utterance to derive an appropriate phonetic baseform. The novelty of the approach is in its employment of an orthography-driven n-best hypothesis and rescoring strategy of the pronunciation alternatives. We make use of decision trees and heuristic tree search to construct and score the n-best hypotheses space. We use acoustic alignment likelihood and phone transition cost to leverage the empirical evidence and phonotactic priors to rescore the hypotheses and refine the baseforms.