Flexible Evidence Model to Reduce Uncertainty Mismatch Between Speech Enhancement and ASR Based on Encoder-Decoder Architecture
Ryu Takeda, Yui Sudo, Kazunori Komatani · 2023
Missing-data automatic speech recognition (ASR) can use the uncertainty of speech enhancement (SE) results to improve ASR performance without joint-training. Such uncertainty is represented by a probabilistic evidence model, whose design affects the performance. Previous evidence models have utilized a specific distribution regardless of the ASR model, and do not consider their robustness, i.e., the ability to recognize both clean and noisy speech (robust ASR). The SE uncertainty depends on the degree of speech cleanness, which causes an uncertainty mismatch between SE and a robust ASR model, resulting in recognition failures. We thus propose a flexible evidence model to reduce the mismatch. Our model features several parameters to control its distribution shape and can reweight the SE uncertainty via an encoder-decoder architecture by automatically selecting parameters in accordance with the ASR scores. Experiments showed that the character error rate improved by more than 3 points over the best baseline even with a robust ASR model under a −5-dB SNR condition.