Integrating evidence over time: A look at conditional models for speech processing.
Eric Fosler‐Lussier, Jeremy Morris, Ilana Heintz, Rohit Prabhavalkar · The Journal of the Acoustical Society of America · 2010
Many acoustic events, particularly those associated with speech events, can be thought of as events in a rich descriptive subspace where the dimensions of the subspace can be thought of as a sort of decomposition of the original event space. In phonetic terms, we can think of how phonological features can be integrated to determine phonetic identity; for auditory scene analysis we can look how features like harmonic energy and cross-channel correlation come together to determine whether a particular frequency corresponds to target speech versus background noise. Some success has been achieved by thinking of these problems as probabilistic detection of acoustic (sub-)events. However, event detectors are typically local in nature and need to be smoothed out by looking at neighboring events in time. This talk describes using conditional random fields models within the automatic speech recognition setting to combine bottom-up speech event detectors. The talk will explore some of the successes and limitations of this log-linear method which integrates local evidence over time sequences.