Identification of Problematic Data Sections and Interpolation of Air Pollution Time-Series
Dave Riegert, David J. Thomson, Aaron Springford · ISEE Conference Abstracts · 2018
Obtaining reliable results from time-series analysis requires complete and error free data and many statistical tools require contiguous data to be usable at all. In the context of measurements from the National Air Pollution Surveillance (NAPS) network we discuss obtaining a data set for use in analysis. We focus on first identifying problematic sections of data and correcting these problems through interpolation, if applicable. A major concern during this process is that because this data is used in decision making it is especially important to avoid introducing new problems while attempting to repair other identified problems.In order to begin addressing any issues present we first must identify that a problem exists, quantify the problem, and flag the points in the data to be corrected. The simplest problem section to identify is one where the data are completely missing. More difficult to identify are sudden or gradual baseline shifts, runs of zero values, runs of saturated values, timing shifts in the data, quantization changes over time, and changes in the autocovariance structure of the data.If a section is to be interpolated we use an interpolation algorithm that is a hybrid Weiner interpolation and consists of an initial interpolation followed by filling each gap individually. This process is iterated until convergence.We discuss: 1) the framework used to allow reproducability of correction routines and the ability to traceback which operations were performed in the case that any questions of validity arise; 2) the approaches used to identify problematic sections of data; and 3) the interpolation algorithm used.