Two Strategies for Text Parsing

Joakim Nivre · 2006

In a previous paper (Nivre, 2005) I have discussed two different notions of parsing that appear in the literature on natural language processing. The first, which I call grammar parsing, is the well-defined parsing problem for formal grammars, familiar from both computer science and computational linguistics; the second, which I call text parsing, is the more open-ended problem of parsing unrestricted text in natural language, which I define as follows: Given a text T = (x1,..., xn) in language L, derive the correct analysis for every sentence xi ∈ T. The main conclusion in Nivre (2005) is that grammar parsing and text parsing are in many ways radically different and therefore require different methods. In this paper, I will concentrate on text parsing and compare two different methodological strategies, which I call the grammar-driven and the data-driven approach (cf. Carroll, 2000). To some extent, these approaches can be seen as complementary, and many existing systems combine elements of both. Nevertheless, from an analytical perspective it may be instructive to contrast the different ways in which they tackle the problems that arise in parsing unrestricted natural language text. 1.1 Grammar-Driven Text Parsing In the grammar-driven approach to text parsing, a formal grammar G is used to define the language L(G) that can be parsed and the class of analyses to be returned for each string in the language. Given my

Read the paper · More papers on PaperTik