Cost-based query optimization in iMeMex

Lukas Blunschi · Repository for Publications and Research Data (ETH Zurich) · 2007

Personal Information Management (PIM) refers to managing all data related to a single individual.In more detail, PIM is: integrating highly heterogeneous data sources, runtime environments which may change dramatically, a query language which is powerful, but still simple to use, semantic integration of all data pertaining to a user, sharing parts of my data with friends, query processing in such a system.PIM is becoming more and more important, but still there is no single system available which is able to handle all challenges.iMeMex [15] is among the first systems aiming to support all of that functionality.In this thesis we enrich iMeMex so that it can adapt to adding or removing data sources and index structures and look at how to process and optimize queries in iMeMex.To solve the former we use OSGi as the basic runtime platform and separate our system into basic building blocks, aka bundles.By doing so not only functionality to access different data sources, but also all indexes, materialized views and replicas become hot-pluggable.As an effect our query processor does not have to be changed when new system functionality is added or removed.To work on the latter we introduce a query processing model that is completely based on query pushdown.Here we profit from having an adaptable system architecture enabling us to treat data sources, indexes and materialized views as low-level query planners.Each of these query planners may decide by itself how much of a query plan it wants to handle and also what cost model it uses.We show how the returned physiological plans can be used in extensible cost-based query optimization.To show the feasability of these techniques, several use cases are studied in detail.

Read the paper · More papers on PaperTik