Verifying Artificial Neural Network Classifier Performance Using Dataset Dissimilarity Measures
Darryl Hond, Hamid Reza Asgari, Daniel Jeffery · 2020
The specification and verification of algorithms is vital for safety-critical autonomous systems which incorporate deep learning elements. In this paper, we introduce a novel measure for quantifying the dissimilarity between the datasets used for training ANN-based image classification algorithms, and the test datasets used for verifying and evaluating classifier performance. The novel dissimilarity measure is defined in terms of each neuron's distribution of output values during training. This measure and its related variants, allow performance metrics, such as accuracy, to be placed into context by characterizing the test datasets employed for classifier evaluation. A system-level requirement could specify the permitted form of the functional relationship between classifier performance and the dissimilarity measures - placing demands on the classifier's capability to generalize to data progressively more distant from the training dataset; such a requirement can be verified by dynamic testing. Empirical results, obtained using publicly available datasets, suggest that the measures have relevance to real-world practice for both quantifying dataset dissimilarity, and specifying and verifying classifier performance.