Machine Learning for Natural Language Processing
Martin Emms · 2011
Over the past decade Machine Learning techniques has become an essentialtool for Natural Language Processing. This introductory course will coverthe basics of Machine Learning and present a selection of widely used al-gorithms, illustrating them with practical applications to Natural LanguageProcessing. The course will start with a survey of the main concepts inMachine Learning, in terms of the main decisions one needs to make whendesigning a Machine Learning application, including: type of training (su-pervised, unsupervised, active learning etc), data representation, choice andrepresentation of target function, choice of learning algorithm. This will befollowed by case studies designed to illustrate practical applications of theseconcepts. Unsupervised learning techniques (clustering) will be describedand illustrated through applications to tasks such as thesaurus induction,document class inference, and term extraction for text classification. Super-vised learning techniques covering symbolic (e.g. decision trees) and non-symbolic approaches (e.g. probabilistic classifiers, instance-based classifiers,support vector machines) will be presented and case studies on text clas-sification and word sense disambiguation analysed in some detail. Finally,we will address the issue of learning target functions which assign struc-tures to linear sequences covering applications of Hidden Markov Modelsto part-of-speech tagging and Probabilistic Grammars to parsing of naturallanguage.