Speech Enhancement Algorithm Analysis for a Reliable Speech Recognition System using Artificial Intelligence Methods
S. Janani, Akhil Hassan G, S Madhankumar, M. Arunkumar · 2023
Speech is the primary means of human communication. Speech has the potential to be a more effective interface than keyboards and pointing devices. A speech interface would support a wide range of useful applications, including phone directory assistance, "hands busy" medical applications, office dictation devices, etc. This has stimulated research on Automatic Speech Recognition (ASR). Speech recognition has grown quickly in recent years thanks to advancements in parameter extraction tools for spoken signals. Automatic speech recognition (ASR) in noisy environments is still a challenge since there are many possible environmental distortions and it is hard to accurately compensate for them. Low recognition performance in noisy environments is mostly caused by mismatches between training and test circumstances. Noise and distortion are the main issues limiting communication systems. Thus, the modeling and removal of distortion and noise effects have been the cornerstones of communications theory and practice as well as signal processing. Noise reduction and distortion removal are important difficulties in voice recognition, picture processing, medical signal processing, radar, sonar, and any other application where the signals cannot be isolated from noise and distortion. There is noise in almost all auditory environments. In applications related to speech, sound recording, telecommunications, speech recognition, and human machine interfaces, the signal of interest typically speech is usually contaminated by noise coming from multiple sources. Today’s speech recognition systems' acoustic modeling components are primarily based on the Hidden Markov Model (HMM).This HMM is well renowned for being a successful model for speech signals and is the most commonly used paradigm for speech recognition. The primary reasons for the reduction in performance are typically the non-native speaker’s lack of fluency and the phonetic discrepancies between the target language and mother tongue. Requiring front-end signal processing is necessary to make voice recognition systems feasible. This is because better signal feature extraction leads to better recognition performance. Voice Activity Detection (VAD) and Speech Enhancement Algorithm (SEA) are used in a preprocessor to improve Recognition Accuracy (RA) in ASR. In order to increase the percentage of ASR’s Recognition Accuracy (RA), this thesis looks at a novel approach for speech enhancement and voice activity detection techniques. A hybrid approach was also proposed to increase the Recognition Accuracy % in different noisy scenarios. In comparison to the proposed VAD and Speech Enhancement algorithms, the recommended hybrid algorithm performs better for variable noise levels at varying Signal to Noise Ratio (SNR) values. Better RA was seen for the proposed hybrid technique in the presence of station noise (86.28%) at different noise levels.