A perceptual representation of sound for source separation.

Daniel P. W. Ellis, Barry L. Vercoe · The Journal of the Acoustical Society of America · 1992

Building machines that emulate the kinds of acoustic information processing that human beings perform effortlessly has proved unexpectedly difficult. The irresistible conclusion is that the human auditory system is extremely sophisticated in its adaptation to real-world sounds and uses an impressive array of features as cues to organization and interpretation. As more of these cues become known through psychoacoustical experiment, it becomes feasible to program computers functionally to mimic human perception of sound. Since the processing is so integrated and the roles of different features only partly understood, the best approach to simulating real listeners (including susceptibility to illusions) is to build as direct an analog of the actual processing chain as can be devised. Concentrating on the task of distinguishing and separating individual superimposed sonic sources, such as a singer and accompaniment, a perceptually sufficient invertible representation has been built that analyzes sound down to a relatively small number of ‘‘atoms.’’ Psychoacoustic rules of stream formation can then be applied to these atoms to simulate many aspects of human source separation. Examples will be shown and played.

Read the paper · More papers on PaperTik