Adaptive data management techniques in modern network infrastructures

Alex Delis, Vassil Kriakov · 2009

We address the general problem of dynamic data management in distributed environments composed of networked clusters of workstations. More concretely, we analyze the problem of handling highly dynamic data sets in two frameworks: (1) indexed multidimensional data sets and (2) distributed high-frequency data streams. Under the first framework, we introduce mechanisms that work off of specialized data structures to allow for dynamic data redistribution in the cluster in such a manner as to capitalize on all available resources while adapting to constantly changing access patterns. In a number of applications, it is not feasible to index and store the underlying data. This calls for an alternative treatment of such continuously updated data under the model of a data stream. We propose a solution for performing sliding window join queries on a set of distributed data streams processed by a cluster of networked workstation. Our method allows for automatic throughput throttling based on resource availability at the cost of approximation errors. Our solution is based on discrete Fourier transform (DFT) representation of the correlations between streams arriving at remote nodes in the system. Furthermore, we provide formulae for computing DFT compression factors in order to achieve information reduction and thus, reduce the cost of inter-node communications. In order to ascertain the effectiveness of our methods, we build prototypes and perform extensive experimentation under high-frequency data updates in a cluster of workstations. Our experiments reveal that the proposed methods scale well in terms of throughput, achieving linear communication complexity under a variety of data distributions. We also investigate the problem of adaptive neighborhood selection in an effort to improve clustering of users in peer-to-peer networks. Our approach is based on the exploitation of both reputation and content similarity. Simulation experiments with the PeerSim engine show improvements in topology, query performance and quality of results compared to currently available Gnutella-like protocols.

Read the paper · More papers on PaperTik