The decodability of sound sources and categories from their acoustic representations as they develop over time

Mattson Ogg, L. Robert Slevc · The Journal of the Acoustical Society of America · 2017

The temporal dynamics of sound impose constraints on how listeners can so rapidly and effectively identify objects in their environment. However, it is unclear what physical features distinguish acoustic sources, and how the roles of those features change as a sound unfolds. Potential differences between the optimal acoustic features for distinguishing sounds and those used by listeners could reveal new insights into how the auditory system functions and which aspects of sound it prioritizes. Thus, we investigated 216 high quality sound tokens decomposed into a set of acoustic features derived from an ERB filter bank in 5 millisecond increments. Support vector machine classifiers were iteratively trained and tested on sound categories (instrument, speech, environmental) and sound sources within those categories (e.g. instruments, speakers) at each time point. Decoding analyses were conducted on acoustic features (e.g. spectral centroid, aperiodicity) and the raw ERB filter output. Categories were decoded better using high level acoustic features, whereas individual exemplars were decoded better from the filter bank output, although these patterns varied over time. By examining how the classifiers weight individual features, these data reveal the relative contributions of specific acoustic features to sound discrimination as a function of time.

Read the paper · More papers on PaperTik