Communication Smoothness Estimation Using F0 Information
Yumi Wakita, Shunpei Matsumoto · 2016
The purpose of this study is to explore a system which provides a topic of discussion for carrying on lively and smooth human-to-human communication. When sensing there has been little progress during the conversation, the system attempts to provide a topic for leading a smoother discussion. To develop a process for estimating conversation smoothness, we confirmed the effectiveness of using fundamental frequency (F0). For each utterance, we calculated the F0’s average values and the F0’s standard deviation (SD) values, and compared these values between “smooth” utterances and “non-smooth” utterances. The differences are shown to be significant when using a t-test, where the confidence level is 95%. The linear discriminant analysis (LDA) were applied to the distribution between the Ave_F0 and the SD for classifying all utterances to “smooth” and “non-smooth”. The classification rates, which mean “% of correctly classified” are over 70% for four speakers. If laughter sounds are included in the utterances, the estimation performance decreased extremely. To keep high performance, it is necessary to except the laughter parts from utterances. We confirmed the Ave_F0 and the SD_F0 value are also effective to detect the laughter sounds. The t-test where the confidence level is 95% was applied on the distribution of the Ave_F0 and the SD_F0 for each speaker. As a result of t-test, the difference between the laughter parts and the speech parts are significant for all four speakers.