Evolving better stoplists for document clustering and web intelligence
Mark P. Sinka, David Corne · 2003
Abstract: Text classification, document clustering and similar document analysis tasks are currently the subject of significant global research, since such areas underpin web intelligence, web mining, search engine design, and so forth. A fundamental tool in such document analysis tasks is a list of so-called ‘stop ’ words,