Development and Evaluation of a Dataset for Evaluating Japanese Speech Style Similarities and Speech Style Embedding Models

Yuki Zenimoto, Shinzan Komata, Ryo Hasegawa, Takehito Utsuro · Transactions of the Japanese Society for Artificial Intelligence · 2025

Dialogue systems are expected to keep its speech style consistency. However, evaluating the similarity ofspeech styles, especially in Japanese, is challenging due to the wide variety of possible styles and the large numberof style-specific vocabulary and expressions. Existing approaches that rely on clustering or manually assigning discretestyle parameters to words cannot fully capture the extensive range of styles observed in real world scenarios.Moreover, classification-based approaches trained on a limited number of known style categories fail to handle unseenstyles. To address these issues, this study proposes a speech style embedding model for Japanese sentences,constructed by fine-tuning a pre-trained BERT model with contrastive learning. Leveraging large-scale data automaticallycollected from utterances inWeb novels, we treat utterance pairs from the same speaker as “positive” examples(similar style) and pairs from different speakers as “negative” examples (dissimilar style). We additionally create andrelease a new human-annotated dataset, JS3 (Japanese Speech Style Similarity), comprising a diverse set of Japaneseutterance pairs labeled with style similarity. Through quantitative evaluation, we demonstrate that our model effectivelydistinguishes not only typical styles such as “ninja” or “princess” but also more nuanced dimensions ofpoliteness. In further analyses, hierarchical clustering (Ward’s method) on the resulting style embeddings revealsdistinct clusters of speech styles, each associated with characteristic vocabulary and word usage patterns. Finally, byexamining the style embeddings of individual speakers, we highlight that even the same character’s speech style canvary substantially depending on conversational partners and surrounding circumstances.

Read the paper · More papers on PaperTik