A Neural-Network Driven Methodology for Anomaly Detection in Apache Spark

Ahmad Alnafessah, Giuliano Casale · 2018

In cloud computing services, multiple tenants sharing the same computing resources can cause performance anomalies, either due to normal or malicious user behaviors. When a Big Data application is deployed over a private or public cloud and does not perform as well as expected, it is however challenging to reliably detect a performance anomaly and minimize its consequences. We here consider in particular Spark-based workloads, in which the analytic operations are applied to a resilient distributed dataset (RDD). We develop a neural network based methodology for anomaly detection based on knowledge of the RDD characteristics. Using experiments against multiple workloads and anomaly types, we show that our method improves over other types of classifiers as well as against black box performance anomaly detection.

Read the paper · More papers on PaperTik