Towards searching the Holy Grail in automatic music and speech processing - examples of the correlation between human expertise and automated classification

Bożena Kostek · 2022

This tutorial lecture is dedicated to addressing the still not fully achieved potential in automatic music and speech processing. There are parallels between music and speech as both are audio signals; however, they differ much in detail. One of such parallels are automatic music transcription and text-to-speech technology; both are very advanced in research and technology but suffer from challenges related to polyphony and differentiated articulation in music, whereas in speech: differentiated accents, L2 speakers' pronunciations, diarization, just to name a few, cause problems. Another challenge concerns affective computation and predicting emotions in music and speech with sufficient accuracy. These challenges lie in the notion that we would like to get better results with machine learning-based classification than human expertise could deliver. In this lecture, first, bases of automatic music and speech processing are to be shortly reviewed. Then, the state-of-the-art in machine learning, including both baseline and deep learning methods, is briefly examined in the context of music and speech processing. Also, advances in available technology related to such issues are presented. Finally, examples of research performed in the discussed areas are shown.

Read the paper · More papers on PaperTik