A Study on Pathology Speech Recognition Based on 2D-Attention-Branchformer
Mo Li, Wei Zhao, Tiantian Yuan · 2024
Communication between speech-impaired and non-disabled people has always been more complex, and the development of artificial intelligence offers the possibility of solving this problem. Speech impairment mainly refers to difficulty in sound formation and aphasia due to brain or peripheral neuropathy; in addition, deafness can also lead to difficulty in sound formation, especially for deaf people who are prelingually deaf. While traditional communication relies heavily on pen on paper and verbal communication, modern technology has enriched communication by providing us with more innovative assistive features, such as speech recognition (ASR). Among these assistive functions, speech recognition technology is essential, as it can capture and recognize the characteristics of various speech signals, opening up a wide range of possibilities for future research and applications. However, speech recognition technology has encountered some difficulties and limitations in recognizing pathological articulations. To better meet the rehabilitation needs of patients, we set up a three-level classification of “light,” “medium,” and “heavy.” In this study, we introduced the 2D-Attention module and integrated the Branchformer technology. First, we utilize the SpecAugment data augmentation technique to mask the acoustic spectrograms with random time and frequency to enhance the diversity of the training data. Next, we utilize the 2D-Attention mechanism to capture better the acoustic spectrogram's time and frequency domain information. The output of 2D-Attention is then used as input to Branchformer for further processing through joint coding. This helps to help pathological speech patients correct their mispronunciations and improve their speech fluency. Experimental results on the speech disorder dataset show that the word error rate of our proposed approach is reduced by 17.7%, 1.4%, 1.6%, and 1.4%, respectively, compared to the mainstream RNN, Transformer, Conformer, and Branchformer models. This approach offers new possibilities for facilitating effective communication between people with speech impairments and others.