Extracting comprehensible models from trained neural networks

Mark W. Craven, JUDE W. SHAVLIK · 1996

for their support and encouragement. Although neural networks have been used to develop highly accurate classifiers in numerous real-world problem domains, the models they learn are notoriously difficult to understand. This thesis investigates the task of extracting comprehensible models from trained neural networks, thereby alleviating this limitation. The primary contribution of the thesis is an algorithm that overcomes the significant limitations of previous methods by taking a novel approach to the task of extracting comprehensible models from trained networks. This algorithm, called Trepan, views the task as an inductive learning problem. Given a trained network, or any other learned model, Trepan uses queries to induce a decision tree that approximates the function represented by the model. Unlike previous work in this area, Trepan is broadly applicable as well as scalable to large networks and problems with high-dimensional input spaces. The thesis presents experiments that evaluate Trepan by applying it to individual networks and to ensembles of neural networks trained in classification, regression, and reinforcement-learning

Read the paper · More papers on PaperTik