Advanced machine learning algorithms for discrete datasets
Sameen Mansha · The University of Queensland · 2020
Despite recent works in the area of machine learning, there remains the need for robust, yet easily usable, methods. In this thesis, we focus on the design, performance, and improvement of well-known clustering and classification algorithms for discrete datasets with application in different domains. In the first section of the thesis, we formulate an optimization problem for clustering interesting itemsets to extract a sparse representation of itemsets and show that their discrete nature makes it NP-hard. An efficient approximation algorithm is presented which greedily solves maximum set cover to reduce overall compression loss. Furthermore, we incorporate our sparse representation algorithm into a layered convolutional model to learn nonredundant dictionary items. Following the intuition of deep learning, our convolutional dictionary learning approach convolves learned dictionary items and discovers statistically dependent patterns using chi-square in a hierarchical fashion; each layer having a more abstract and compressed dictionary than the previous. In the second section for fairness aware classification, we utilize reject option in different classifiers, a general decision-theoretic framework for handling instances whose labels are uncertain, for modelling and controlling discriminatory decisions. Specifically, this framework permits a formal treatment of the intuition that instances close to the decision boundary are more likely to be discriminated in a dataset. We propose three different solutions for discrimination-aware classification problems. The first solution invokes probabilistic rejection in single or multiple probabilistic classifiers while the second solution relies upon ensemble rejection in classifier ensembles. The third solution integrates one of the first two solutions with situation testing which is a procedure commonly used in the court of law. We evaluate our proposed clustering and discrimination-aware classification solutions on relevant benchmark real-world datasets and compare their performance with previously proposed state of the art approaches. The results demonstrate the superiority of our solutions in terms of performance and flexibility of applicability.