Multimodal Voice Activity Prediction: Turn-taking Events Detection in Expert-Novice Conversation

Kazuyo Onishi, Hiroki Tanaka, Satoshi Nakamura · 2023

Predicting the timing of utterances in dyadic conversations is essential for achieving natural interactions between humans and virtual agents. Since the former often use non-verbal cues to adjust the order of their speech, this study proposes a multimodal model incorporating non-verbal features using a Transformer-based voice activity prediction model. First, in line with previous research, we reproduced a baseline model that utilized audio features (audio waveform, voice activity frame, and voice activity history) as inputs. To this baseline model, we added non-verbal features: gaze direction, action units, head pose, and articular points. We compared our multimodal model with the baseline model to investigate the impact of non-verbal cues on voice activity prediction. We utilized a dyadic expert-novice conversation dataset and evaluated the average outcomes across ten model trainings. Results revealed that our proposed models with all the features improved the accuracy of the next speaker prediction by 2.3% and back-channeled prediction by 1.8% (p-value < 0.025). In particular, action units may contribute significantly to the turn-shift and back-channeled predictions. This study demonstrates that including non-verbal features in Transformer-based turn-taking models enhances the efficacy of models for predicting voice activity in dyadic conversations.

Read the paper · More papers on PaperTik