Correcting, Rescoring and Matching: An N-best List Selection Framework for Speech Recognition

Chin-Hung Kuo, Kuan-Yu Chen · 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC) · 2022

In recent years, automatic speech recognition (ASR) has been widely used in various scenarios, and it is usually the first step in many applications. Therefore, more and more studies concentrate on enhancing the recognition results. Among them, N -best reranking and error correction models are two active research subjects. Various models have been proposed and demonstrated their success. However, as the N -best reranking models aim to select the best hypothesis from a set of candidates, their performance upper bound is limited by the given set of hypotheses. The error correction models detect and correct recognition errors so as to provide better results, but they usually perform the process on the highest-scored hypothesis only. Therefore, the information embedded in other candidates is ignored. Besides, we note that almost all of the N -best reranking and error correction models consider the acoustic information implicitly, indirectly, or even omitted. In order to mitigate these flaws, we propose an N -best list selection framework, which consists of a text correction module, a text rescoring module, and a text-speech matching module, for speech recognition. Based on the proposed framework, a set of corrected hypotheses can be deduced, and then the text rescoring module is introduced to accurately rescore them. In addition, the text-speech matching module is employed to calculate the alignment score between each hypothesis and its own speech. The proposed framework is evaluated on the AISHELL-l dataset, and the experimental results reveal that the proposed framework can deliver over 30 % character error reduction rates compared to the baseline systems.

Read the paper · More papers on PaperTik