Effects of Multiple Japanese Datasets for Training Voice Activity Projection Models

Yuki Sato, Yuya Chiba, Ryuichiro Higashinaka · 2024

Voice Activity Projection (VAP) is crucial in dialogue systems for facilitating natural turn-taking. So far, models for predicting speech activity in spoken conversations and models incorporating multimodal information have been reported. However, these models primarily handle English audio data, and there is limited research using other languages, including Japanese. In this study, we utilized the Transformer-based VAP model proposed by Ekstedt and Skantze and trained it on multiple Japanese datasets. Then, we compared performance differences depending on the dataset. The results demonstrate the utility of the model trained on Japanese data for predicting Japanese speech activity. We also obtained insight into how the types of dialogue in the datasets affect the performance difference.

Read the paper · More papers on PaperTik