A Novel Approach to Overcome Data Scarcity Problem for Highly Inflecting Languages
K. V. N. Sunitha, A. Sharada · 2010
This paper proposes a model for lexicon which can be effectively used in language modelling and other NLP related tasks for agglutinative, highly inflecting and compounding languages. Due to huge amount of distinct word forms, the traditional methods based on full words are not very effective and it is not straight forward to train efficient language models with good coverage of the language. The main contribution of the paper is the proposal of new data structure and an algorithm that is applied over the data structure. This approach greatly reduces the corpus size thereby making the research work in the field easy. The use of the new data structure fastens the search process.