Using Graphs in Construction of a Lemmatization Model for Turkish
Enis Arslan, Umut Orhan · Çanakkale Onsekiz Mart University AVESIS · 2017
In this paper, a lemmatization framework which uses the capability of a graph database is presented. Asintroduced in the previous research, using Finite State Machines (FSM) in lemmatization of Turkish words isapplicable when an affix-stripping method is preferred. These studies present results for limited datasets andcan be modelled for larger and actual data environments. To ensure a living up-to-date system we proposea dynamic lemmatization model which feeds up a static graph database model with new words by using amophological function to validate the graph relations. This function is developed on Finite State Machines(FSM) which encodes Turkish grammar and affixes. Good results on the framework can lead to discovery ofout-of-vocabulary (OOV) words and disambiguation of the ambiguous ones