Information Retrieval Using Concept Lattices
Arvind Kumar Muthukrishnan · OhioLink ETD Center (Ohio Library and Information Network) · 2006
Information retrieval concerns the problem of extracting useful and relevant information.Information retrieval algorithms are used almost in all departments e.g.education, sales, product reviews etc.There are many general and domain specific search engines like Google (general search engine), ACM (domain specific search engine) which implement good information retrieval algorithms.The general search engines flood the user with a lot of results and many irrelevant data are also present.Due to this more and more domain specific information retrieval systems are being developed.These systems retrieve better results than general search engines but they are also plagued by the same problems.Currently there are different information retrieval methods like Clustering, Vector Space model, Latent Semantic Indexing etc. being used.Most of these systems use keyword based retrieval which checks for the presence of an entered keyword in documents and ranks them according to their frequency.These systems won't retrieve documents where the concept represented by the keyword is present but the keyword itself is absent.In this thesis we discuss about an information retrieval system which employs concept based retrieval, i.e. this system retrieves relevant documents based on the concept represented by the keyword rather than just the presence of the keyword.The data set we use is the documents/papers and the attributes associated with them.We collect data from four different fields namely 'Data Mining', 'Dynamic Programming', 'Graph Algorithms' and 'Networks'.In this system manual indexing is used to capture all the important concepts/attributes as determined by the domain experts.The documents are covered by attributes from three different perspectives namely 'Structural Information', 'Content Information' and 'Publication Information' providing a better insight.We use concept lattices to represent data and the lattice structure can be queried for interesting results.In this system only relevant documents are retrieved so the user isn't flooded with a lot of data, which makes it easy for the user to browse through the results to find the desired document.Also, complex queries involving multiple fields/collections can be executed and paths between two separate searches can be displayed.viii