Real-time Anomaly Prediction at Scale using Anomaly Detection Augmented with Regression
Kunal Banerjee, Binay Gupta, Meet Maheshwari, Lalitdutt Parsai, Geet Vudata, Soumik Dasgupta, Anirban Chatterjee · 2024
Anomalies are rare untoward events which lead to system failures, business interruptions and bad customer experiences. While anomaly detection (AD) tries to identify these anomalies as quickly as possible so that these incidents can be resolved at the earliest, anomaly prediction (AP) has a more ambitious goal of foretelling the anomalies before these even occur so that the damaging effects of anomalies can be avoided altogether. In short, AD is a reactive technique, whereas AP is a proactive technique. However, as with any proactive technique, some of the predictions may be wrong – too many false predictions may hinder the adoption of an AP solution due to alert fatigue. Therefore, it is important that the AP solution not only identifies almost all the anomalies in advance (i.e., has high recall) but does so with minimal false predictions (i.e., has high precision). In our work, we showcase how an AP solution can be built using an existing AD solution and augmenting it with a regressor. We carried out extensive experiments on the SKAB public dataset with multiple combinations of AD classifiers with regressors, and found that the AP solution improves the AD solution irrespective of the choice of the classifier or the regressor. Thus, this work sheds light on how any existing AD solution may be extended to perform anomaly predictions as well. Furthermore, we apply this technique at an industry-scale for predicting real-time anomalies in Kafka channels – our solution has been able to predict incidents in advance by up to 55 minutes while reducing alert fatigue by more than 62%.