Transformation-based error-driven learning and natural language processing: a case study in part-of-speech tagging

Eric Brill · 1995

Recently, there has been a rebirth of empiricism in the eld of natural language processing. Manual encoding of linguistic information is being challenged by automated corpus-based learning as a method ofproviding a natural language processing system with linguistic knowledge. Although corpus-based approaches have been successful in many different areas of natural language processing, it is often the case that these methods capture the linguistic information they are modelling indirectly in large opaque tables of statistics. This can make it di cult to analyze, understand and improve the ability of these approaches to model underlying linguistic behavior. In this paper, we will describe a simple rule-based approach to automated learning of linguistic knowledge. This approach has been shown for a number of tasks to capture information in a clearer and more direct fashion without a compromise in performance. We present a detailed case study of this learning method applied topart of speech tagging. 1.

Read the paper · More papers on PaperTik