DOCUMENT RETRIEVAL USING CLUSTERING

B. Sivaranjani · International Journal of Research in Engineering and Technology · 2015

The exponential growth of knowledge in the World Wide Web, has understood the need to develop economical and effective ways for organizing relevant contents.In the field of web computing, document clustering plays a vital role and plays an interesting and challenging problem.Document clustering is mainly used for grouping the similar documents in the search engine.The web also has rich and dynamic collection of hyperlink information.The retrieval of relevant document from the internet is the complicated task.Based on the user's query the document will be retrieved from the various databases to give relevant information and additional information for the given query.The documents are already clustered based on keyword extraction and stored in the database.The probabilistic relational approach for web document clustering is to find the relation between two linked pages and to define a relational clustering algorithm based on probabilistic graph representation.In document clustering, both content information and hyperlink structure of web page are considered and document is viewed as a semantic units.It also provides additional information to the user.

Read the paper · More papers on PaperTik