Error Correction for Speech Recognition Systems Using Large Language Model Reasoning Capabilities

Sun Lina, Konstantin A. Aksyonov · 2024

In this research, we have meticulously engineered an Automatic Speech Recognition system through a postprocessing error correction approach. Within the scope of this endeavor, we have integrated existing Large Language Models as the foundational framework. It is widely acknowledged that Large Language Models exhibit commendable performance in text processing tasks. Nonetheless, the recalibration of a new large-scale model tailored to specific task demands entails a significant expenditure of computational resources. The crux of this paper is the strategic utilization of pre-existing large models, coupled with the exploitation of the inherent characteristics of LLMs, such as contextual inferential capabilities, probabilistic reasoning, and internal linguistic rule processing, to refine and correct the ASR system’s output. We have deployed the Chat3 Large Language Model as the pivotal component of the ASR system’s error correction module, employing an “ensemble model fusion technique” to bolster the system’s performance. This methodology epitomizes an academically rigorous and resource-efficient approach to augmenting the efficacy of ASR systems without necessitating the extensive training of new models.

Read the paper · More papers on PaperTik