Unveiling Data Preprocessing Patterns in Computational Notebooks

Valentina Golendukhina, Michael Felderer · 2024

Data preprocessing, which includes data integration, cleaning, and transformation, is often a time and effort-intensive step due to its fundamental importance. This crucial phase is integral for ensuring the quality and suitability of data for sub-sequent stages, such as feature engineering and model training in Machine Learning-enabled and data-driven systems. This paper provides an extensive overview of data preprocessing functions in Python and examines their application and prevalence in computational notebooks by analyzing 149,048 computational notebooks collected from Kaggle. Despite the crucial role played by data preprocessing in model performance, our results expose a significant lack of emphasis on data preprocessing activities in the examined notebooks. Notably, users holding the highest rankings tend to skip data preprocessing steps and focus on model-related activities. Although other users exhibit more frequent incorporation of data preprocessing methods, the overall prevalence remains relatively limited. We discovered that data preparation practices such as missing values are present in 20 % to 60 % of the notebooks depending on the competition, whereas outliers handling is only present in less than 20% of the analyzed scripts. The most frequently and consistently applied practices are the data transformation methods.

Read the paper · More papers on PaperTik