On statistical methods in natural language processing

Joakim Nivre · DSpace repository (University of Tartu) · 2001

What is a statistical method and how can it be used in natural language processing (NLP)? In this paper, we start from a definition of NLP as concerned with the design and implementation of effective natural language input and output components for computational systems. We distinguish three kinds of methods that are relevant to this enterprise: application methods, acquisition methods, and evaluation methods. Using examples from the current literature, we show that all three kinds of methods may be statistical in the sense that they involve the notion of probability or other concepts from statistical theory. Furthermore, we show that these statistical methods are often combined with traditional linguistic rules and representations. In view of these facts, we argue that the apparent dichotomy between "rule-based" and "statistical" methods is an over-simplification at best.

Read the paper · More papers on PaperTik