Single channel speech enhancement using principal component analysis and MDL subspace selection

Rolf Vetter, NATHALIE VIRAG, Philippe Renevey, Jean-Marc Vésin · 1999

Landmark based speech processing is a component of Lexical Access From Features (LAFF), a novel paradigm for feature based speech recognition. Detection and classification of landmarks is a crucial first step in a LAFF system. Vowel landmarks are detected using an existing syllabic segmentation algorithm with several novel extensions that incorporate durational information, absolute energy level, and F1 track information. The detector is scored against the TIMIT database, using a novel algorithm to convert the segmental transcriptions to a landmark representation for scoring. Previous experiments have validated the predictions of acoustic theory, specifically the presence of F1 peak in vowels, and demonstrated that amplitude peak in a fixed frequency band is practically as good as a formant tracker. Substantial improvements in performance are achieved by optimizing the fixed frequency band to the F1 range, and use of a trainable neural network to combine the multiple acoustic cues. The neural network can also be used to generate confidence scores for detected landmarks, which provide vital information for later stages of processing.

Read the paper · More papers on PaperTik