Information Integration Across Heterogeneous Sources: Where Do We Stand and How to Proceed?
Aditya Telang, Sharma Chakravarthy, Yan Qun Huang · 2008
Today, information integration has assumed a completely different, complex connotation than what it used to be. The advent of the Internet, the proliferation of information sources on the surface Web as well as the deep Web, the presence of structured, semistructured, and unstructured data- all have added new dimensions to the problem of information integration as known earlier. From the time of distributed databases leading to heterogeneous, federated, and multi-databases, retrieval and integration of information from heterogeneous sources has been an important and complex problem. Currently, the problem is even more complicated as repositories exist in various formats (HTML, XML, spatial data sources to name a few) and schemas, and both the content and the structure of the data within them are changing autonomously. As the number of repositories/sources will continue to increase in an uncontrolled manner, there is no other option but to find extensible techniques for answering a complex search/query whose (partial) answers have to be retrieved and integrated from multiple sources. In this survey paper, we identify the set of challenges that need to be addressed for this form of heterogeneous information integration, and compare the current state-of-the-art as to how they fare. We then propose a framework with functional components – termed InfoMosaic, that aims to address some of these important challenges, and briefly elaborate on the data and control flow involved in answering a complex query/search.