Linguistic Features

Björn Wolfgang Schuller, Anton M. Batliner · 2013

A multiplicity of methods exist when dealing with linguistic analysis, some of which including deeper linguistic analysis. Thus, this chapter presents a subset of the predominant approaches. It shows different approaches that were mostly introduced for the processing of text comprised of words; yet, they can be transferred to any domain dealing with sequences of symbols. The chapter talks about ‘words’ consisting of ‘characters’ representing the basic string units of analysis, and use ‘speech’ and ‘text’ in the sense of ‘audio with symbolic content’ and ‘symbolic content’. After text pre-processing, linguistic descriptors are extracted, most commonly in the form of feature vectors with real-valued components (vector space modelling). The dimensions of these vectors (the vocabulary) usually correspond to either words (‘bag-of-words’), sequences of words (‘bag-of-N-grams’), or sequences of characters (‘bag-of-character-N-grams’).

Read the paper · More papers on PaperTik