Low dimensional synthetic data generation for improving data driven prognostic models
Tony Lindgren, Olof Steinert · 2022
Data driven prognostic models are becoming more prevalent in many areas, ranging from heavy trucks to gas turbines. One aspect of certain prognostic models is the need for labeled failures, which then can be used as positive examples, when modelling the prognostic problem. Unfortunately, standard algorithms for creating prognostic models can suffer when labeled data is unbalanced, w.r.t. class distribution, leading to prognostic models with poor performance. In this paper we present a methodology for creating synthetic data that can be used to augment the underrepresented class and hence dramatically increase performance of the data driven predictive model. In our study we utilize data collected from heavy trucks and focus on predicting failure of one engine component that is crucial for the operation of heavy trucks. We examine different way of generating synthetic examples in a low dimensional setting, it is found that three methods out of the six methods studied does not improve performance compared to using only the original data. The other three methods based on interpolation is superior to only using the original data, with SMOTE outperforming the two other interpolation methods. SMOTE lowers the estimated cost on test data, compared to using a model trained on the original data set only, with 67%.