The decision trees and the optimization of resources in Big Data solutions
Youssf En-nattouh, Khalid El Fahssi, Ali Yahyaouy, Jamal Riffi, Hamid Tairi · 2020
every day, we see that a quantitative explosion of digital data has forced researchers to find new strategies to collect, store, analyze and visualize data. In the context of storage and processing of a large massive amount of data we find a lack of powerful tools to master and control them. Also, during the process of executing tasks in real time in clustered IT platforms, we encounter the problem of optimizing parallel tasks. So we will propose in this article a method based on the algorithm of decision trees as an interpretable machine learning algorithm which can allow us to evaluate the impact of certain characteristics on the variable of the task execution time. This decision tree algorithm is useful and it helps us understand how we can optimize the different parameters that affect workloads in clustered applications. We can thus optimize the number of tasks in Big Data clustered applications without failure and performance degradation.