Acoustic Features
Björn Wolfgang Schuller, Anton M. Batliner · 2013
This chapter describes the extraction of acoustic features from the speech signal. It starts with digitalisation of the speech signal, pre-processing and enhancement, short time analysis for time-domain and frequency-domain low-level descriptor (LLD) extraction. The choice of LLDs is based on those most frequently found in the field of computational paralinguistics. For many applications a continuous audio stream can be directly analysed frame-by-frame. In other applications, however, chunking is required for ‘supra-segmental’ analysis. More complex solutions are based on multi-dimensional feature information and machine learning. These approaches can be trained well to the signals of interest and therefore usually allow better results — but at the cost of greater effort. Pitch detection algorithms (PDAs) in the time domain analyse the speech signal period by period and determine the periods' boundaries.