Properties of a GP active learning framework for streaming data with class imbalance
Sara Khanchi, Malcolm Iain Heywood, Nur Zincir-Heywood · Proceedings of the Genetic and Evolutionary Computation Conference · 2017
Active learning algorithms attempt to interactively develop a subset of data from which fitness evaluation is performed. Moreover, the distribution of labeled content within the data subset may adapt over time as genetic programming (GP) individuals improve. The basic goal is therefore to identify the most meaningful subset of data to improve the current model. Under a streaming data context additional challenges exist relative to the non-streaming scenario: non-stationary processes, partial observability anytime operation. This means that it is not possible to guarantee that the content of the data subset even provides exemplars for each class that could appear in the stream (i.e., different classes appear/disappear at different parts of the stream). With this in mind, an investigation is performed into the impact of adopting different policies for controlling the development of data subset content. To do so, a generic framework is defined in terms of sampling and archiving policies. The resulting evaluation under several large multi-class datasets with class imbalance indicates that adopting random sampling with a biased archiving policy is sufficient for evolving GP classifiers that match or better the current state-of-the-art, particularly when detecting minor classes.