Supporting the optimisation of distributed data mining by predicting application run times
Shonali Priyadarsini Krishnaswamy, Seng W. Loke, Arkady Zaslavsky · International Conference on Enterprise Information Systems · 2003
There is an emerging interest in optimisation strategies for distributed data mining in order to improve response time. Optimisation techniques operate by first identifying factors that affect the performance in distributed data mining, computing/assigning a to those factors for alternate scenarios or strategies and then choosing a strategy that involves the least cost. In this paper we propose the use of application run time estimation as solution to estimating the cost of performing a data mining task in different distributed locations. A priori knowledge of the response time provides a sound basis for optimisation strategies, particularly if there are accurate techniques to obtain such knowledge. In this paper we present a novel rough sets based technique for predicting the run times of applications. We also present experimental validation of the prediction accuracy of this technique for estimating the run times of data mining tasks.