Near Language Identification Using NooJ

Božo Bekavac, Kristina Kocijan, Marko Tadić · Repozitorij Filozofskog fakulteta u Zagrebu' at University of Zagreb (University of Zagreb) · 2015

In this work we took a linguistic knowledge aware approach tailored for a specific pair of languages. We use NooJ as a core part of a system designed for automatic identification of near languages, Croatian and Serbian in particular. We use several levels of NooJ processing capabilities. First, we apply specially designed lexical transducers for the detection of the typical morphological spots in language. Then we apply the syntactic grammars for the detection of verb da verb syntagmas, characteristic for Serbian language. Finally, we measure discrepancies between texts provided by text processing. The output is generated according to predefined voting principle using AutoHotkey program. Our results show high F1 measures for language identification of Croatian and Serbian texts.

Read the paper · More papers on PaperTik