Speaker Pseudonymization for Japanese Speech Using Duration Embeddings

Aoi Ito, Katunobu Itou · 2024

In order to safely expand the use of speech data, research is being conducted on speaker pseudonymization as a method to protect the privacy of speech data. Speaker pseudonymization is a technology that processes speech data to preserve information that can be used for subsequent tasks such as speech recognition while concealing the speaker. In machine learning-based methods, the mainstream approach is to process x-vectors created from MFCCs to change information about speech quality. In this study, we focus on duration information of phoneme or mora as an element that can identify individuals other than speech quality, and aim to expand the scope of pseudonymization of speech data by processing duration as a feature that represents individuality. The proposed pseudonymization based on duration and its conversion reduced the accuracy rate of speaker recognition models by 22% and expanded the scope of pseudonymization of speech data.

Read the paper · More papers on PaperTik