Enhancing Drift Detection and Model Uncertainty Handling in Imbalanced Streaming Data using Autoencoder-based Approach

Shubhangi S. Suryawanshi, Anurag Goswami, Pramod Ravindra Patil · 2023

In today’s digital era, numerous applications are generating data in the form of data stream. The data streams are a continuous massive amount of data generated in real-time. Handling the data streams in real-time environments is a critical challenge. The data stream classification becomes more challenging when the statistical properties of the features (changes in the underlying data distribution) or their relationship with the target label are changing and the phenomenon is known as concept drift. When concept drift occurs, a machine learning model trained on previous data may no longer predict or generalise to new incoming data accurately. This can lead to a decrease in performance, poor accuracy, and unreliable predictions. In this paper, we propose an approach for drift detection using an Auto-encoder-based model that incorporates KL divergence and reconstruction loss, along with class imbalance handling and integrating the Monte Carlo Dropout. The Auto-encoder learns to capture the inherent patterns and regularities, facilitating drift detection by computing the reconstruction loss. KL divergence is an additional loss term, which encourages the Auto-encoder to learn a latent space that is sensitive to changes in the data distribution. Borderline-Synthetic Minority Oversampling Technique (Borderline-SMOTE) is used to address the challenge of class imbalance in drift detection, which generates synthetic minority samples. Integrated the Monte Carlo Dropout, a probabilistic regularization technique, to enhance the model’s uncertainty estimation. The proposed model was evaluated on real-world datasets. The results showed that the proposed model outperformed the other state-of-the-art methods.

Read the paper · More papers on PaperTik