Topic Models for Audio Mixture Analysis

Paris Smaragdis, Madhusudana V. S. Shashanka · 2009

In contrast to the time-domain waveform, time-frequency representations explicitly represent the time-varying frequency content of a sound and effectively visualize the signal’s activity at any timefrequency bin. These transforms are often complex-valued and include well known tools such as the short-time Fourier transform, constant-Q transforms, wavelets, etc. However, because our hearing system is more sensitive to the relative energy between different frequencies, for most practical applications we study the modulus of these transforms and discard the phase which is useful only in special cases. These kinds of representations are essentially counting the number of time-frequency acoustic quanta that collectively make up complex sound scenes, similar to how we count words that make up documents. With this representation we can use the analogy of bag of frequencies, which we describe later in this document.

Read the paper · More papers on PaperTik