Implementation of query optimization techniques in distributed environment through genetic algorithm

Mishra Sambit Kumar, Srikanta Pattnaik, Dulu Patnaik · 2014

In distributed databases, usually the cost of optimization increases while consulting the underlying data sources in the query optimization process. The query optimizer may handle the data sources and the optimization techniques adopted are required to be implemented to all relevant cost information with minimal communication. It is understood that query optimization issues depend upon relations i.e. cardinality, size of a tuple and fraction of tuples participating in a join with another relation. Also it is seen that the query optimization issues depend upon attribute comprising with cardinality of domain, actual number of distinct values and Common assumptions. The query optimization task involves with selecting the plans and sub plans, evaluating the size of plans and sub plans along with cost of plans and sub plans. In many cases the query execution cost is measured in conceptual units. But In a distributed database, the query execution costs must be divided into multiple dimensions. In this paper it is intended to discuss the approaches of query optimization and enumerate the query optimization techniques in distributed environment. As it has been seen that selecting the optimal execution strategy for a query is NP-hard in the number of relations, genetic algorithm may be applied in this case to evaluate plan select value, CPU time, cost of query plans and sub plans including I/O cost. The purpose of doing so is to measure the cost of data sets while dealing with large data sets in the distributed environment.

Read the paper · More papers on PaperTik