Well Said: An Analysis of the Speech Characteristics in the LibriSpeech Corpus

Diptasree Debnath, Helard Becerra Martinez, Andrew Hines · 2023

Recent trends in speech quality research have shown the effectiveness models trained with large datasets of unlabelled natural speech. The LibriSpeech corpus contains approximately 1000 hours of English language read audiobooks and has been a popular dataset for many data driven speech technology models from automatic speech recognition to accent recognition and speech quality prediction. While the curators of LibriSpeech balanced the dataset to ensure there was a balance of speakers and content without overlaps, speech characteristics were not a design consideration. In this paper we use six algorithms to analyse speech for pitch, intensity and rate across the seven subsets of the LibriSpeech corpus. We find a good distribution between speakers and within some speakers for the characteristics tested. We show that speech characteristics are well balanced across subsets. We conclude that this validation makes LibriSpeech corpus to assist in the development of an objective model for synthetic speech quality prediction.

Read the paper · More papers on PaperTik