Auditory scene analysis as Bayesian inference in sound source models

Maddie Cusimano, Luke Hewitt, Josh Tenenbaum, Josh H. McDermott · 2018 Conference on Cognitive Computational Neuroscience · 2018

Inferring individual sound sources from the mixture of soundwaves that enters our ear is a central problem in auditory perception, termed auditory scene analysis (ASA).The study of ASA has uncovered a diverse set of illusions that suggest general principles underlying perceptual organization.However, most explanations for these illusions remain intuitive (or are narrowly focused), without formal models that predict perceived sound sources from the acoustic waveform.Whether ASA phenomena can be explained by a small set of principles is therefore unclear.We present a Bayesian model based on representations of simple acoustic sources, for which a deep neural network is used to guide Markov chain Monte Carlo inference.Given a sound waveform, our system infers the number of sources present, parameters defining each source, and the sound produced by each source.This model qualitatively accounts for perceptual judgments on a variety of classic ASA illusions, and can in some cases infer perceptually valid sources from simple audio recordings.

Read the paper · More papers on PaperTik