A Comparative Analysis of Unsupervised Learning Techniques for Anomaly Detection in Railway Systems
Macílio da Silva Ferreira, Lucio Flavio Vismari, Paulo Sérgio Cugnasca, Jorge Rady de Almeida, João Batista Camargo, Guilherme Miranda Kallemback · 2019
Anomaly detection is a fundamental Data Mining activity. It is largely applied to identify data with potential to contaminate or bias the results of an analysis. On the other hand, anomalies can also indicate faults in real-world systems. In the railway context, anomalous data can occur due to faults on its assets (elements). Thus, it is possible to use the anomaly detection task both to monitor the system health and to avoid future consequences, especially accidents. Due to the immense volume of data contained in a railway system, unsupervised learning techniques are commonly used for this purpose. However, no studies were found that evaluate the efficiency of these techniques in anomaly detection, mainly applied to large real datasets from real railway systems. This paper thus presents a comparative analysis among unsupervised learning techniques applied to anomaly detection in the behavior of track switches, an availability and safety critical asset in a railway system. The K-means, Self-Organizing Maps, and Auto-encoders techniques were applied to real dataset from a railway in Brazil to detect anomalies in their track switches. A comparison of the technique performances in detecting anomalies in a real dataset were developed, resulting in a ranking. These results can be used as a basis for developing predictive methods for fault detection and diagnosis in railway systems.