Novel Active Learning Framework for Anomaly Detection in Aviation with Expert in the Loop
Milad Memarzadeh, Bryan L. Matthews, Thomas Templin, Aida Sharif Rohani, Daniel Weckler · AIAA SCITECH 2022 Forum · 2022
View Video Presentation: https://doi.org/10.2514/6.2022-2542.vid Anomaly detection in commercial aviation is an extremely challenging yet crucial task. Accurately detecting operationally significant anomalies in operational data can save civilian lives and/or result in significant savings in aircraft/maintenance cost. The current practice uses manually tuned rule-based mechanisms to flag exceedances from pre-defined safety boundaries. However, this system cannot identify unknown risks and emerging vulnerabilities. Recently, innovative approaches based on data science and machine learning have been utilized to automate anomaly detection. However, there are limits to their applicability in the field of aviation due to several fundamental challenges: (1) Properly reviewed data is scarce in aviation and, as a result, supervised learning approaches cannot reach optimal performance. (2) Operationally significant anomalies do not coincide with statistically significant ones and, as a result, unsupervised learning approaches fail to provide reliable and robust performance. In this paper, we propose SALAD, a Semi-supervised Active Learning framework for Anomaly Detection, which detects operationally significant anomalies in flight operational quality assurance data. The developed framework works with vast amounts of unlabeled data as well as a small quantity of labeled data reviewed by subject matter experts (SMEs) to reliably identify safety anomalies in flight operations. Moreover, the model’s active learning strategy allows it to detect unknown risks and anomalies that might emerge in the system earlier with the help of SME reviews. We validate performance of the SALAD framework and its components with a real-world case study of multi-class anomaly detection during the approach to landing of commercial aircraft. We specifically show that the proposed framework reaches reliable anomaly-detection accuracy when only one percent of the data is reviewed and labeled and can identify unknown anomalies effectively.