Time-scaling of speech using independent subspace analysis

Rangarao Muralishankar, A. G. Ramakrishnan, Lakshmish Kaushik · 2004

We propose a new technique for modifying the time-scale of speech using Independent Subspace Analysis (ISA). To carry out ISA, the single channel mixture signal is converted to a time-frequency rep-resentation such as spectrogram. Here, the spectrogram is gener-ated by taking Hartley or Wavelet transform on overlapped frames of speech. We do dimensionality reduction of the autocorrelated original spectrogram using singular value decomposition. Then, we use Independent component analysis to get unmixing matrix using JadeICA algorithm [5]. It is then assumed that the over-all spectrogram results from the superposition of a number of un-known statistically independent spectrograms. By using unmix-ing matrix, independent sources such as temporal amplitude en-velopes and frequency weights can be extracted from the spec-trogram. Time-scaling of speech is carried out by resampling the independent temporal amplitude envelopes. We then obtain time-scaled independent spectrograms after multiplying the independent frequency weights with time-scaled temporal amplitude envelopes. Summing all these independent spectrograms and taking inverse Hartely or wavelet transform of the sum spectrogram to recon-struct and overlap-add the reconstructed time-domain signal to get the time-scaled speech. The quality of the time-scaled speech has been analyzed using Modified Bark Spectral Distortion(MBSD) [6]. From the MBSD score, one can infer that the time-scaled sig-nal is less distorted. 1.

Read the paper · More papers on PaperTik