Robust speech recognition by properly utilizing reliable frames and segments in corrupted signals
Yi Chen, Chia-yu Wan, Lin-shan Lee · 2007
In this paper, we propose a new approach to detecting and utilizing reliable frames and segments in corrupted signals for robust speech recognition. Novel approaches to estimating an energy-based measure and a harmonicity measure for each frame are developed. SNR-dependent GMM Classifiers are then trained, together with a Reliable Frame Selection and Clustering module and a Reliable Segment Identification module, to detect the most reliable frames in an utterance. These reliable frames and segments thus obtained can be properly used in both front-end feature enhancement and back-end Viterbi decoding. In the extensive experiments reported here, very significant improvements in recognition accuracies were obtained with the proposed approaches for all types of noise and all SNR values defined in the Aurora 2 database.