Data Quality Visual Analysis (DQVA) A tool to process and pinspot raw data irregularities
Célio Carvalho, Rui Silva Moreira, José Manuel Torres · 2021
This project proposes a machine learning (ML) pipeline for inferring office employee's well-being, from heterogeneous sources of contextual data (cf. physiological, social and workplace environment), which brings several demanding issues. In this paper we focus specifically in raw data collection problems and pre-processing challenges. To start with, context data was collected in real environments, during weeks, in several office organizations and involving employees along theirs daily working routines. Moreover, data collection resort to a wide range of sources (e.g. sensors, questionnaires, apps, etc.) that were subject to potential interferences and noisy conditions. Given the influence of data quality in ML algorithms results and considering the number of instruments used, it was essential to implement a pre-processing stage to automate and improve the quality of collected data. Hence, the usefulness of the proposed DQVA tool, which computes several common statistical measures and provides also graphical and tabular visual insights about the data. For example, it allows to: i) compare data sources from different participants and organizations, on a per sensor/data source basis (through data tables, data distribution histograms, and visualizations); iii) check and pinspot the existence of outliers; iv) visually spot signal gaps; etc. Therefore, we argue that the proposed DQVA tool allows to evaluate, per sensor and per individual, raw data quality, on the integration stage of our classification pipeline. It proved to be an agile, useful and simple to re-use tool for detecting raw data irregularities, thus increasing data quality assurances for the next steps of our classification pipeline.