Test model for stop word removal of devnagari text documents based on finite automata
Anjusha Pimpalshende, Anjali Mahajan · 2017 IEEE International Conference on Power, Control, Signals and Instrumentation Engineering (ICPCSI) · 2017
In the digital era ample data is available on internet. Processing of unstructured data is todays need. In IR (information retrieval systems), Web Mining, Artificial Intelligence, Natural Language Processing, Text Summarization, Text and Data Analytic systems, optimization of text data becomes very important. One of the preprocessing step is stop word removal. Some extremely common words which would appear to be of little value in helping select documents matching a user need are excluded. These words are called stop words. Stop words list has been developed for languages like English, Chinese, Arabic, Hindi, etc. A large number of available works on stop word removal techniques are based on manual stop word lists. An efficient stop word removal technique is required. In this paper, we are proposing a stop word removal algorithm for Devanagari Languages. Which is using the concept of a Finite Automata (DFA). Then pattern matching technique is applied and the matched patterns, which is a stop word, is removed from the document. Previous methods are time consuming, as searching process takes a long time. In comparison of that, our algorithm gives better result in less time.