Unsupervised Machine Learning: The Apriori Algorithm
Jan Novotny, Paul Bilokon, Aris Galiotos, Frédéric Délèze · 2019
The unsupervised machine learning techniques can be easily solved for a very small number of dimensions as the joint probability function can be directly estimated. On the other hand, this is not possible in large dimensions, and various approximations are used. The favourite choices are for example variations to Gaussian mixtures. The authors note that the dimensionality of the feature vector is usually much larger than in the case of supervised machine learning problems. Further, researchers are trying to find patterns in high-dimensional data with various forms of clustering analysis, principal component analysis, independent component analysis, or self-organising maps. This chapter focuses in detail on another class of algorithms known as association rules. In particular, it discusses and implements the Apriori algorithm. The chapter encourages the reader to stretch the limits of the algorithm and investigates its performance on much bigger sets, preferably not using in-memory tables.