A First Step Towards Parsing of Assamese Text
Navanath Saharia, Utpal Sharma, Jayantee Kalita · 2011
Assamese is a relatively free word order, morphologically rich and agglutinative language and has a strong case marking system stronger than other Indic languages such as Hindi and Bengali. Parsing a free word order language is still an open problem, though many different approaches have been proposed for this. This paper presents an introduction to the practical analysis of Assamese sentences from a computational perspective rather than from linguistics perspective. We discuss some salient features of Assamese syntax and the issues that simple syntactic frameworks cannot tackle. Keywords-Assamese, Indic, Parsing, Free word order. I. INTRODUCTION Like some other Indo-Iranian languages (a branch of Indo- European language group) such as Hindi, Bengali (from Indic group), Tamil (from Dardic group), Assamese is a morphologically rich, free word order language. Apart from possessing all characteristics of a free word order language, Assamese has some additional characteristics which make parsing a more difficult job. For example one or more than one suffixes are added with all relational constituents. Research on parsing model for Assamese language is purely a new field. Our literature survey reveals that there is no annotated work on Assamese till now. In the next section we will present a brief overview of different parsing techniques. In section III we discuss related works. Section IV contains a brief relevant linguistic background of Assamese language. In section V we discuss our approach we want to report in this paper. Section VI conclude this paper.