Benchmarking morphological analyzers for the Hungarian language
Gábor Szabó, László Kovács · Az Eszterházy Károly Tanárképző Főiskola tudományos közleményei. Tanulmányok a matematikai tudományok köréből/Az Eszterházy Károly Főiskola tudományos közleményei. Tanulmányok a matematikai tudományok köréből/Annales mathematicae et informaticae · 2018
In this paper we evaluate, compare and benchmark the four most widely used and most advanced morphological analyzers for the Hungarian language, namely Hunmorph-Ocamorph, Hunmorph-Foma, Humor and Hunspell.The main goal of the current research is to define objective metrics while comparing these tools.The novelty of this paper is the fact that the analyzers are compared based on their annotation token systems instead of their lemmatization features.The proposed metrics for the comparison are the following: how different their annotation token systems are, how many words are recognized by the different analyzers and how many words are there whose morphological structure is equivalent using a well-defined mapping among the annotation token systems.For each of these metrics, we define the concept of similarity and distance.For the evaluation we use a unique Hungarian corpus that we generated in an automated way from Hungarian free texts, as well as a novel automated token mapping generation algorithm.According to our experimental results, Hunmorph-Ocamorph gives the best results.Hunmorph-Foma is very close to it, but sometimes returns an invalid lemma.Humor is the third best analyzer, while Hunspell is far worse than the other three tools.