IMPROVING RETRIEVAL ACCURACY IN WEB DATABASES USING INTRA-TABLE AND INTER-TABLE DEPENDENCIES

Ravi Gummadi · 2009

This thesis deals with query answering over the Web databases. Since the Web databases are independently and autonomously populated the Web users, it leads to different problems. One such problem is missing key information and rendering the direct join infeasible. Primary key-foreign key (PK-FK) information lies at the heart of traditional databases and assists in joining tables. In the recent years, increasing amounts of data are populated by lay users into autonomous Web databases such as Google Base and Amazon SimpleDB. This has lead to an absence of any centralized control over the data being populated. Issues such as missing data, imprecise queries and missing PK-FK information began creeping into Web databases. In this thesis, a system to deal with the problem of missing PK-FK is described. The SMARTINT system contains three important modules Source Selection, Query Processing and Learning. The key idea underlying the framework is to exploit the mined attribute dependencies present in the data and use them to select a tree of tables which is subsequently expanded to form the result set. The performance and the accuracy of SMARTINT has been thoroughly evaluated over test data crawled from Google Base. The precision and recall of the results given by SMARTINT are significantly higher compared to direct join and single table approaches which validates the proposed solution. SMARTINT showed an average of 55% higher accuracy (F-measure) than the other two approaches. It also showed the same amount of improvement in accuracy over all the possible joins. Its learning module showed over 80% improvement in the execution time over ‘state of the art’ approaches.

Read the paper · More papers on PaperTik