A Study on Document Clustering

Baby Maheshwari, Ahmed Abdul, Moiz Qyser, Subhash Chandra · 2013

Information Retrieval (IR) is a large and growing field within Natural Language Processing (NLP). The search engine is the most well-known (and perhaps still the only really useful) application. Search engines like Google1 and AltaVista2 are used by many people on a daily basis. There are several other applications within IR. Among them we consider text clustering in particular. Document clustering (also referred to as Text clustering) is closely related to concept of data clustering. Document clustering is a more specific technique for unsupervised document organization, automatic topic extraction and fast information retrieval or filtering. For example, a web search engine often returns thousands of pages in response to a broad query, making it difficult for users to browse or to identify relevant information. Clustering methods can be used to automatically group the retrieved documents into a list of meaningful categories, as is achieved by Enterprise Search engines such as Northern Light and Vivisimo. The main idea is to find which documents have many words in common, and place the documents with the most words in common into the same groups. It is done without using any predifined categories. Text clustering can for instance be applied to the documents retrieved by a search engine, so that they can be presented in groups according to content.

Read the paper · More papers on PaperTik