Natural language processing

Gobinda G Chowdhury · Annual Review of Information Science and Technology · 2003

Previous ARIST chapters (Haas, 1996;Warner, 1987) described a number of theoretical developments that have influenced research in NLP.The most recent theoretical developments can be grouped into four classes: (i) statistical and corpus-based methods in NLP, (ii) recent efforts to use WordNet for NLP research, (iii) the resurgence of interest in finite-state and other computationally lean approaches to NLP, and (iv) the initiation of collaborative projects to create large grammar and NLP tools.Statistical methods are used in NLP for a number of purposes, e.g., for word sense disambiguation, for generating grammars and parsing, for determining stylistic evidences of authors and speakers, and so on.Charniak (1995) points out that 90% accuracy can be obtained in assigning part-of-speech tag to a word by applying simple statistical measures.Jelinek (1999) is a widely cited source on the use of statistical methods in NLP, especially in speech processing.Rosenfield (2000) reviews statistical language models for speech processing and argues for a Bayesian approach to the integration of linguistic theories of data.Mihalcea & Moldovan (1999) mention that although thus far statistical approaches have been considered the best for word sense disambiguation, they are useful only in a small set of texts.They propose the use of WordNet to improve the results of statistical analyses of natural language texts.WordNet is an online lexical reference system developed at Princeton University.This is an excellent NLP tool containing English nouns, verbs, adjectives and adverbs organized into synonym sets, each representing one underlying lexical concept.Details of WordNet is available in Fellbaum ( 1998) and on the web (http://www.cogsci.princeton.edu/~wn/).WordNet is now used in a number of NLP research and applications.One of the major applications of WordNet in NLP has been in Europe with the formation EuroWordNet in 1996.EuroWordNet is a multilingual database with WordNets for several European languages including Dutch, Italian, Spanish, German, French, Czech and Estonian, structured in the same way as the WordNet for English (http://www.hum.uva.nl/~ewn/).The finite-state automation is the mathematical device used to implement regular expressions -the standard notation for characterizing text sequences.Variations of automata such as finite-state transducers, Hidden Markov Models, and n-gram grammars are important components of speech recognition and speech synthesis, spell-checking, and information extraction which are the important applications of NLP.Different applications of the Finite State methods in NLP have been discussed by Jurafsky & Martin (2000), Kornai (1999) and Roche & Shabes (1997).The work of NLP researchers has been greatly facilitated by the availability of large-scale grammar for parsing and generation.Researchers can get access to large-scale grammars and tools through several websites, for example Lingo

Read the paper · More papers on PaperTik