Speaker Identification Method Using Facial Image and Voice
Etsuro Nakamura, Yoichi Kageyama, Motonari Shirasu · 2020 IEEE 2nd Global Conference on Life Sciences and Technologies (LifeTech) · 2020
The task of creating meeting minutes can be substantially simplified with the use of an automatic system, and the efficiency of subsequent meetings and operations can also be improved. In particular, the function of automatically assigning speakers to minutes is crucial. Two methods of speaker identification have been widely utilized in practice; a method of assigning speakers to cameras, and a method of registering speech in advance. However, in each case, it is necessary to optimally arrange the equipment. In this paper, we propose a method for discriminating between speakers using an omnidirectional camera and a microphone. This method does not require multiple devices or the pre-registration of audio. In this process, image and voice data are used for speaker identification. In speaker discrimination experiments, 83.8 % of the speakers were successfully discriminated.