A Structured Approach for Building Assamese Corpus: Insights, Applications and Challenges
Shikhar Kr. Sarma, Himadri Bharali, Ambeswar Gogoi, Ratul Deka, Anup Kr. Barman · International Conference on Computational Linguistics · 2012
To study about various naturally occurring phenomenons on natural language text, a well structured text corpus is very much essential. The quality and structure of a corpus can directly influence on performance of various Natural Language Processing applications. Assamese is one of the major Indian languages used by the people of north east India. Language technology development works in Assamese language have been started at various levels, and research and development works started demanding a structured and well covered Assamese Corpus in UNICODE format. Here we present various issues and problems related to building an Assamese text corpus. We review our experience with constructing one such corpus including about 1.5 million words of Assamese language. It will provide a significant effort by serving as an important research tool for language and NLP researchers.