Algorithm A for distributed data Classification

Evans Teiko Tetteh, Beata Marta Zielosko · Procedia Computer Science · 2024

Knowledge discovery is one of the key areas in predictive data mining tasks. Performing Classification tasks on a single source of data using a decision tree algorithm is a relatively straightforward process. However, the complication arises when we have distributed sources of data that yield sets of decision trees. Classifier ensembles are often created, and a decision is assigned to a new object based on a certain voting strategy. The article proposes a different approach to creating a rule-based classifier. Using a set of decision trees, a global model of decision rules is induced. It contains rules which are true for the maximum number of trees from a set of decision trees. This model is verified by data corresponding to distributed local data sources. The bootstrapping technique was used to obtain distributed data sources. Pruning of decision trees was applied to improve the accuracy of the rule-based classifier. The conducted experiments confirm the validity of using algorithm A for the learning of decision rules from a set of decision trees.

Read the paper · More papers on PaperTik