Multi-resolution auditory scene analysis: robust speech recognition using pattern-matching from a noisy signal
Sue Harding, Georg F Meyer · 2003
Abstract Unlike automatic speech recognition systems, humans canunderstand speech when other competing sounds are presentAlthough the theory of auditory scene analysis (ASA)may helpto explain this ability, some perceptual experiments show fu-sion of the speech signal under circumstances in which ASAprinciples might be expected to cause segregation. We pro-pose a model of multi-resolution ASA that uses both high- andlow- resolution representations of the auditory signal in paral-lel in order to resolve this conflict. The use of parallel repre-sentationsreduces variabilityforpattern-matching whileretain-ing the ability to identify and segregate low-level features ofthe signal. An important feature of the model is the assump-tion that features of the auditory signal are fused together un-less there is good reason to segregate them. Speech is recog-nised by matching the low-resolution representation to previ-ously learned speech templates without prior segregation of thesignal into separate perceptual streams; this contrasts with theappr oach gener a l l y used by comput at i onal m odels of A S A . Wedescribe an implementation of the multi-resolution model, us-ing hidden Markov models, that illustrates the feasibility ofthis approach and achieves much higher identification perfor-mance than standard techniques used for computer recognitionof speech mixed with other sounds.