Investigating the Effectiveness of Speaker Embeddings for Shout Intensity Prediction

Takahiro Fukumori, Taito Ishida, Yoichi Yamashita · 2023

The automatic detection of shouted speeches has attracted much research attention as a core technology of audio surveillance systems. A common strategy in the past literature has been to train a binary classifier using labels of shouted or normal speeches. Although it is known that the acoustic properties of shouted speech usually differ among speakers, especially male and female groups, the conventional methods did not pay attention to encoding such personal and gender-related style information. There are recent findings that speaker embeddings, which are produced by a model trained for speaker identification, can improve other tasks such as speech emotion classification. Thus, this paper investigates the effectiveness of such speaker embeddings for a shouted speech detection problem. Specifically, we verify whether x-vector embeddings can work as effective auxiliary features for the target problem compared with a simple gender label of an input speech whose ground truth is generally unavailable in real situations. Our experiments on predicting the shout intensity beyond the traditional binary classification demonstrated the x-vector embeddings achieved performance improvement over the single use of speech features.

Read the paper · More papers on PaperTik