Extending a Japanese Speech−to−Gesture Dataset Towards Building a Pedagogical Agent for Second Language Learning
Dai Hasegawa, Hiroshi ECHIZENYA · Institutional Repositories DataBase (IRDB) · 2021
We created a Japanese speech-to-gesture dataset where we recorded 298 mins (trimmed into 210 mins) of speech audio data and the motion capture data of the accompanying gestures.Our aim was to tackle speech-to-gesture generation by using a data-driven approach on the dataset.Our first attempt of the speech-to-gesture generation partially succeeded.However, to improve data-driven gesture generation, we will need more informative the dataset.In this paper, to enrich the dataset, we annotated seven gesture phases (rest, preparation, pre-stroke hold, stroke, post-stroke hold, retraction) to 240 sentences out of 1047 sentences of our dataset.However, the annotation process needs efforts.Thus, to extend the annotations to all over the dataset, we tested a gesture phase estimation by using Bi-directional Long-Short Term Memory (Bi-directional LSTM).The results show that overall accuracy of seven gesture phase estimation was 0.61 in precision, 0.61 in recall, and 0.61 in f-value.The most successful phase in the estimation was rest phase scored 0.91 in precision, 0.94 in recall and 0.92 in f-value.