Design and optimization of an autonomous feature selection pipeline for high dimensional, heterogeneous feature spaces

Bernhard Schlegel, Bernhard Sick · 2016

The growing complexity and diversity of vehicle systems makes it increasingly hard to identify and resolve the root cause of an unplanned maintenance session at hand in a dealer workshop. This especially holds for workshop staff with limited qualifications and experience. By today, the workshop staff is supported by expert-based systems, where knowledge was manually generated, i. e., by formalizing human knowledge using rules. Due to the above mentioned reasons, this approach is coming to its limits while the potential of car- and workshop-data available today remains unused. A highly autonomous, machine learning based approach seems promising. To unleash its full potential, the selection of relevant features from the multi-thousand dimensional feature space is indispensable. This feature selection needs to take the automotive requirements into account. These include, but are not limited to, sparse data, the requirement that the selected features are interpretable, and the aim that the trained models save warranty costs. This article presents and evaluates a hands-free feature selection pipeline, paving the way for automotive model building. The pipeline consists of three layers: a feature preparation layer, a filter layer that significantly reduces the feature space at low computational cost based on entropy and statistical measures, and a wrapper layer that selects the final feature set for training based on simple models. Finally, highly performant logistic regression models have been trained to generate the metrics used for evaluation.

Read the paper · More papers on PaperTik