Effective Speech Data Augmentation Method To Improve Customer Service Representative Speech Recognition System Performance

Huiyong Bak, Changhyeon Jeong · 2024

In this paper, we propose two data augmentation methods that effectively improve the performance of the customer service representative (CSR) speech recognition system. Building an CSR speech recognition system includes high transcription costs for speech data. Additionally, for use in training, long form CSR speech data is split into one sentence. The performance of the CSR speech recognition system is low because the split data does not contain conversational features between the CSR and the customer. The data augmentation method proposed in this paper uses untranscribed CSR speech data by transcribing it with WhisperX. Additionally, the two CSR speech data divided into one sentence are combined into speech data in the form of a conversation between two people. To verify the proposed data augmentation method, two CSR speech data from AI-HUB were used. As a result of the verification, it was confirmed that the performance of the CSR speech recognition system was improved when the proposed speech data augmentation method was applied.

Read the paper · More papers on PaperTik