Efficient data scheduling for real-time large-scale data-intensive distributed applications

Mohammed Eltayeb · OhioLink ETD Center (Ohio Library and Information Network) · 2004

Data staging is a method for organizing data transfer in heterogeneous distributed computing environments.Real-time large-scale data intensive applications are distributed applications that frequently generate, utilize and access large-scale data sets with constraints for completion times.These applications are currently emerging in many areas of science and engineering.The data-intensive operations performed by these applications -including transfer of large data sets (order of 100's of Gigabyte transfer), can quickly consume network and compute resources and hence incur degraded overall performance for the applications in these distributed networked environments.This research proposes several solutions that enable efficient data staging to solve this vital problem in heterogeneous distributed computing.Our solutions, mainly, maximize the satisfiability of the applications by designing efficient data scheduling heuristics for such systems.Two optimization models for this maximization were proposed: one model for optimizing the overall satisfiability and the other for the optimizing autonomous applications.In this research we introduce and develop deterministic solutions for the problem.We propose three main algorithms: Two for dynamic setting and one for static.Our main static data scheduling heuristic, called Concurrent Scheduling over Extended Partial Path (CS/EPP) heuristic, is developed based on assumed data staging model for optimizing ACKNOWLEDGMENT I wish to thank my adviser, Professor Fusun Ozguner, for intellectual support, encouragement, and enthusiasm which made this thesis possible, and for her patience in correcting both my stylistic and scientific errors.

Read the paper · More papers on PaperTik