A Transfer Learning Approach for Music-driven 3D Conducting Motion Generation with Limited Data
J.S. Oh, Jinwoo Jeong, Young Ho Chai · 2024
Generating motions based on audio using deep learning has been studied steadily. However, previous research has mainly focused on speech-driven 3D gesture generation and music-driven 3D dance motion generation. We aim to generate 3D motions for specific scenarios, such as conducting. To address the challenge of lacking existing training datasets, we constructed a multi-modal 3D conducting motion dataset, which containing 1.43 hours and is, to our knowledge, a small-scale dataset. Furthermore, we propose a novel approach that uses transfer learning with a model pre-trained on a speech gesture dataset to generate 3D conducting motions. We evaluate the generated motions both with and without transfer learning, using quantitative and qualitative metrics. Our results show that the proposed method improves performance in both aspects compared to the baseline without transfer learning.