An automated domain specific stop word generation method for natural language text classification
Hakan Ayral, Sırma Yavuz · 2011
We propose an automated method for generating domain specific stop words to improve classification of natural language content. Also we implemented a Bayesian natural language classifier working on web pages, which is based on maximum a posteriori probability estimation of keyword distributions using bag-of-words model to test the generated stop words. We investigated the distribution of stop-word lists generated by our model and compared their contents against a generic stop-word list for English language. We also show that the document coverage rank and topic coverage rank of words belonging to natural language corpora follow Zipf's law, just like the word frequency rank is known to follow.