A Dictionary Data Processing Environment and Its Application in Algorithmic Processing of Pali Dictionary Data for Future NLP Tasks
Jürgen Knauth, David Alfter · 2014
This paper presents a highly flexible infrastructur e for processing digitized dictionaries and that can be used to build NLP tools in the future. This infrastructure is especially suitable for low resource languages where some digitized information is available but not (yet) suitable for algorithmic use. It allows researchers to do at least some processing in an algorithmic way using the full power of the C# programming language, reducing the effort of manual editing of the data. To test this in practice, the paper de scribes the processing steps taken by making use of this infrastructure in order to identify wor d classes and cross references in the dictionary of Pali in the context of the SeNeReKo p roject. We also conduct an experiment to make use of this data and show the importance of th e dictionary. This paper presents the experiences and results of the selected approach.