Error detection and correction in annotated corpora
Markus Dickinson · OhioLink ETD Center (Ohio Library and Information Network) · 2005
Building on work showing the harmfulness of annotation errors for both the training and evaluation of natural language processing technologies, this thesis develops a method for detecting and correcting errors in corpora with linguistic annotation.The so-called variation n-gram method relies on the recurrence of identical strings with varying annotation to find erroneous mark-up.We show that the method is applicable for varying complexities of annotation.The method is most readily applied to positional annotation, such as part-of-speech annotation, but can be extended to structural annotation, both for tree structuresas with syntactic annotation-and for graph structures-as with syntactic annotation allowing discontinuous constituents, or crossing branches.Furthermore, we demonstrate that the notion of variation for detecting errors is a powerful one, by searching for grammar rules in a treebank which have the same daughters but different mothers.We also show that such errors impact the effectiveness of a grammar induction algorithm and subsequent parsing.After detecting errors in the different corpora, we turn to correcting such errors, through the use of more general classification techniques.Our results indicate that the particular classification algorithm is less important than understanding the nature of the errors and altering the classifiers to deal with these errors.With such alterations, we can automatically correct errors with 85% accuracy.By sorting the errors, we can I would like to thank my adviser Detmar Meurers for his insight into this thesis and for his support and encouragement during my entire time in graduate school.He has taught me how to find my way in computational linguistic research and how to successfully collaborate.I look forward to even more collaboration in the future.I would also like to thank Chris Brew for many discussions on issues both great and small related to this thesis and related to almost any aspect of computational linguistics.He has always been able to provide insightful questions and encourage my research to go in new directions.Bob Levine has also been of enormous help in ensuring that my time in the PhD program was well-spent.He has provided long discussions on syntactic issues and has provided detailed, constructive, and useful feedback on topics outside his specialty, and for that I am both impressed and extremely grateful.All of my professors at OSU have throughout my graduate student career afforded more many opportunities for exploration and discussion and have never let me forget that linguistics is an essential part of computational linguistics.The OSU computational linguistics discussion group at, CLippers, has been invaluable in providing feedback on many different aspects of this thesis.I also thank the audiences at EACL-03, TLT-03, MCLC-04, MCLC-05, the NODALIDA-05 special v session on treebanks, and ACL-05, where portions of this work have been presented before.I also thank the OSU Department of Linguistics and the Edward J. Ray Travel Award committee for the financial support that allowed me to travel to these conferences; and the OSU