COMPARISON OF A SYNTAX-BASED AND A KNOWLEDGE-POOR PRONOUN RESOLUTION SYSTEMS FOR TURKISH

Dilek Küçük, Meltem Turhan Yöndem, Yılmaz Kılıçaslan · 2007

This study presents a comparison of two pronoun resolution systems developed for identifying the antecedents of third-person personal pronouns referring to proper person names in Turkish texts of any genre. The first one is a syntax-based system which employs a simple syntax-based heuristic rule along with simple semantic and syntactic constraints and some preferences, whereas the second one is a knowledge-poor system using limited syntactic knowledge through language-specific constraints and preferences. The implemented systems were run on the same test corpora, whereby manifestations of each approach on Turkish were observed and some future directions to pursue are addressed in this paper with a helping hand from the analysis of these observations. Introduction The main goal of this study is to compare the implementations of two different pronoun resolution algorithms resting on observations on their performance on the same test corpora. The first algorithm is an adaptation of Hobbs' naive approach to Turkish, which can be counted as an example of the traditional linguistics method. The other algorithm is a knowledge-poor one deliberately limiting the use of syntactic, semantic and discourse knowledge in the resolution process, and observe their performance on the same test corpora. We expect to find significant differences not only in performance but also in the types of errors the two algorithms usually make. We also briefly consider whether extensions are possible to improve the performance of these algorithms. The main contribution in this experimental work is the use of the same corpora when evaluating these two algorithms, which is essential for a detailed comparison of error characteristics of the algorithms so that objective conclusions can be drawn to improve the accuracy and performance of these pronoun resolution systems for Turkish. The Syntax-based Pronoun Resolution System for Turkish (1) Hobbs' naive approach (2), which the first algorithm is based on, has attracted considerable attention in the research community as a syntax-based algorithm and is still one of the most successful algorithms: recent comparisons show that it is still on a par with the vast majority of modern resolution systems. The naive algorithm has been reformulated so that it can be applied to Turkish. Then, this reformulated version has been improved by using mainly syntactic information with simple semantic and syntactic constraints, which are derived from Binding Theory (3), as well as some preferences for identifying antecedents of third-person pronouns in Turkish. The Knowledge-poor Pronoun Resolution System for Turkish (4) The knowledge-poor approach, which the second algorithm employs, uses limited syntactic knowledge to resolve pronouns. It makes use of language-specific constraints and preferences in the resolution process where the scope of the system covers third- person personal and reflexive pronouns referring to proper names. This system is the first fully specified knowledge-poor computational framework for pronoun resolution in Turkish, a language which has structural properties different from the languages for which knowledge-poor systems have been developed.

Read the paper · More papers on PaperTik