REMOVING REDUNDANCY FROM SOME COMMON REPRESENTATIONS OF SPEECH

Ladan Baghai-Ravary, SW BEET, Mohammad Osman Tokhi · 2024

This paper shows how a nan-stationary vector: predictor can be used to identify redundancy in various common forms of speech data.A number of different forms of data are used: some producing spectrogram-like representations.while others are rarely displayed mhimlly in the literature, and so many researchers are not familiar with the structure they exhibit (or the fact that they exhibit any signi cant structure at all).The method used here is known as ow-based prediction (FBF) [i].The prediction takes the farm of a standard vector linear predictor [2], but with a sparse, mevat'ying.prediction matrix, which is updated over a very short time scale.This makes it eminently suitable for modelling speech dynamics, since large changes in, for example, formant trajectories, can occur over a very small number of analysis frames.

Read the paper · More papers on PaperTik