Speech/Music Classification of Short Audio Segments
Toni Hirvonen · 2014
Research on speech/music classification of digital audio has been both popular in academia, and increasingly utilized in industry. Most of the usual methods use carefully hand-crafted features with Gaussian Mixture Models. To get best performance, some of the features necessitate a long latency due to look ahead, or/and a long onset error. This paper aims to have a different approach to the problem by exploring some of the latest trends in machine learning that have resulted in improvements in other fields. Specifically, it is shown that we can achieve comparable performance by only analyzing segments in the order of tens of milliseconds without the use of following or previous audio. This is done by using a method that allows automatic generation of arbitrarily many features from preprocessed spectrograms.