A trainable method for pronominal anaphora resolution using shallow information.

Michael D. Paul, Eiichiro Sumita · Journal of Natural Language Processing · 2001

We propose a corpus-based approach to anaphora resolution of Japanese pronouns combining a machine learning method and statistical information. First, a decision tree trained on an annotated corpus determines the coreference relation of a given anaphor and antecedent candidates and is utilized as a filter in order to reduce the number of potential candidates. In the second step, preference selection is achieved by taking into account the frequency information of coreferential and non-referential pairs tagged in the training corpus as well as distance and counting features within the current discourse.

Read the paper · More papers on PaperTik