Applying RNNs Architecture by Jointly Learning Segmentation and Stemming for Myanmar Language
Yadanar Oo, Khin Mar Soe · 2019
Due to the powerful development of internet use, the amount of unstructured Myanmar text data has increased excessive. Stemming has been widely used in a variety of search engine to increase the retrieval accuracy. Stemming is a method that reduces morphology similar variant of word into a single term called stems or roots. Stemming also influence in accuracy of text categorization, Information Retrieval and text summarization etc. Many word stemmers are available for the major languages, but they are not existing for Myanmar. Word segmentation for Myanmar Language, like for most Asian Languages, is an important task and extensively studied sequence labeling problem. There is no space between words and segmentation is essential pre-processing requirement for many natural language processing applications. Segmentation error would cause translation mistakes directly. This approach proposes different types of recurrent neural networks and different layers that jointly learn segmentation boundaries and stemming. Joint word segmentation and stemming of this research is aiming to support Information Retrieval and Myanmar natural language processing applications.