Comparison of Static and Time-Sequential Features in Automatic Fluency Detection of Spontaneous Speech

Huaijin Deng, Takehito Utsuro, Akio Kobayashi, Hiromitsu Nishizaki · 2021

There have been lots of previous studies on disfluency detection. However, most of them focus on lexical cues, and little emphasis is placed on how diverse acoustic features contribute to improving the performance. We describe a framework for automatic fluency evaluation of spontaneous speech. We investigate not only lexical features extracted from transcription, but also consider time-sequential and static acoustic features from audio data, including energy and voicing-related features. This work tries to reveal how diverse acoustic features contribute to the performance of speech fluency evaluation. The proposed framework was evaluated with the Corpus of Spontaneous Japanese using LSTMs and DNN architectures. Evaluation results showed that when detecting fluent speech, combining four time-sequential acoustic features with static lexical features achieved the best performance. When detecting disfluent speech, on the other hand, static jitter/shimmer features helped to improve the precision at relatively high lower bounds.

Read the paper · More papers on PaperTik