A Linear-Clustering algorithm for controlling quality of large scale water-level data in Thailand
Nuttapon Pattanavijit, Peerapon Vateekul, Kanoksri Sarinnapakorn · 2015
Hydro and Agro Informatics Institute (HAII) has installed more than 800 telemetry stations across Thailand to collect water level data for operation tasks and researches, e.g., flooding prevention system. To have an accurate result, it is crucial to control the quality of data by detecting and filtering out anomalies. In our previous work, a data quality management system to capture various types of errors was proposed. However, the algorithms to detect outliers and missing patterns are based on DBSCAN, which requires complicated implementation and excessive computational cost. In this paper, we present a novel clustering algorithm specially designed for water-level data called “Linear Clustering. ” Compared to DBSCAN, it is not only much easier to develop, but it also requires less computational time without losing any detection accuracies. An analysis of the runtime showed that the proposed algorithm requires linear time. Experiments were conducted on large scale water-level data. For outlier detection, the new method took only 3 seconds on 30,000 records of data, while the previous work took 261 seconds. For missing pattern detection, although there is no difference in runtime, Linear Clustering's code is uncomplicated, and therefore it requires less developing time.