Digit Classification System for Normal and Pathological Speech

Y. A. Goutham, T. S. Himasagar, G S Likhith, Nagendra Gowda, Veena Karjigi, Chandrashekar H M · 2024

Pathological speech recognition is the specialized field within speech recognition that focuses on understanding and transcribing speech that deviates from typical patterns due to various medical conditions. This task is essential for developing effective clinical applications, such as diagnostic tools, and assistive technologies that aid individuals with communication difficulties. Signal representations are crucial because they transform the raw speech waveform into a format that machine learning models can easily analyze. A suitable signal representation should clearly emphasize the unique characteristics of pathological speech while being resistant to the noise and other distortions often present in this type of data. This study demonstrates the impact of signal representation by developing a digit recognition system and initially using a short-time Fourier transform (STFT) to convert speech into spectrograms. The system performs well with normal speech data, but its accuracy decreases when tested with pathological speech. This decline is likely due to the limited amount of pathological speech data available for training the model. To address this issue, future work will focus on data augmentation, employing various time-frequency representations to increase the dataset’s size and capture the diverse nature of pathological speech.

Read the paper · More papers on PaperTik