On the Rule-Based Parsing of Czech
Kiril Ribarov · 2002
This paper presents various attempts to accustom the Rule-based Approach (RBA) as originally introduced by Eric Brill in 1993 (Brill 1993a), on the problem of parsing of Czech (a highly inflective language) within a dependency based syntactic framework (Sgall et al 1986). It is experimentally supported in this paper that the modification of RBA for the stated aim is neither simple nor straight-forward for all attempts fail to produce such a set of rules that parses Czech with a comparable success rate to the one obtained for English or the one obtained for the currently best parser (statistical) for Czech (Collins et al. 1999). At the beginning of this article the parsing problem restricted to the need of this work is specified. Further, the results of the first attempts to modify the rules for RBA parsing of Czech is presented, the influence of the cardinality of the tag set size of the parsing success rates is studied and a second attempt at a wider modification of the rule templates is provided. The article contains the basic accompanying experimental results and aims to be a solid background and a starting point for any adaptation of the rule-based approach for the purposes of parsing of an inflective language within a dependency based framework. 1. On the understanding of parsing