SOUND: Sanity Checking of Pipelines for Uncertain and Sparse Data Series

Hermann Stolte, Iftach Sadeh, Elisa Pueschel, Avigdor Gal, Matthias Weidlich · 2025

The analysis of data series forms the basis of decision-making in various domains, so that it is essential to ensure data validity. Yet, current solutions for sanity checking of processing pipelines, such as GX, TFDV, Pandera or Deequ, fall short in accounting for data quality issues. In particular, irregular cadences, sparsity and value uncertainty limit the applicability of sanity checking and pose risks of false conclusions. In this paper, we present Sound to enable sanity checking of pipelines in the presence of typical quality issues in data series. In particular, Sound evaluates a set of sanity constraints that formalize validity expectations on the data, while incorporating data quality issues, i.e., uncertainty of individual data points and sparsity in a whole data series. To this end, it defines a statistical framework for constraint checking that is based on adaptive resampling and Bayesian hypothesis testing, minimizing computational costs while ensuring accurate results. If a constraint violation has been identified, Sound also includes drill-down strategies to guide users in the identification of the root cause of the violation. We demonstrate the feasibility and utility of Sound by applying it for pipelines developed in the domains of smart grid monitoring and astrophysics.

Read the paper · More papers on PaperTik