Methods for noise reduction in a legacy speech corpus

Lisa Lipani, Yuanming Shi, Joshua McNeill, Margaret E. L. Renwick · The Journal of the Acoustical Society of America · 2019

The Digital Archive of Southern Speech is an audio corpus featuring interviews conducted from 1968 to 1983, with speech from 30 female and 34 male southern speakers, totaling 372 h of audio data. However, automated analysis of this data has been made difficult by background noise in this legacy corpus, originally recorded on reel-to-reel tape and later digitized to .wav format. In this paper, we compare the effect of noise reduction techniques on the acoustic signal and evaluate their effect on acoustic speech data. We use Praat to detect the quietest silence (measured using root-mean-square amplitude) in each sound file. The audio data contain silences that lack background noise (e.g., due to anonymization procedures), but these are excluded from selection. Each quietest silence is used to create a “noise profile,” which is removed from the audio using a scripted noise removal procedure in both Audacity and SoX. The success of each procedure is assessed with three-way comparisons of the amplitude of silences in uncleaned and cleaned audio, and vowel plots made with formant values extracted from uncleaned and cleaned audio. It is predicted that successful noise reduction will reduce the number of outliers occurring in F1, F2 space.

Read the paper · More papers on PaperTik