Deducing linguistic structure from the statistics of large corpora

Eric Brill, David M. Magerman, Mitchell P. Marcus, Beatrice Santorini · 1990

Within the last two years, approaches using both stochastic and symbolic techniques have proved adequate to deduce lexical ambiguity resolution rules with less than 3-4% error rate, when trained on moderate sized (500K word) corpora of English text (e.g. Church, 1988; Hindle, 1989). The success of these techniques suggests that much of the grammatical structure of language may be derived automatically through distributional analysis, an approach attempted and abandoned in the 1950s.

Read the paper · More papers on PaperTik