‘Prodata’: A Python Library for Simplifying Manual Data Preprocessing

Poorva Nande, Supriya Kelkar · 2025

Data preprocessing is a crucial phase in the data science and machine learning pipeline, often demanding significant time and expertise. This step is vital for enhancing data quality by eliminating noise and inconsistencies, which in turn improves the performance and reliability of models. By streamlining the dataset, effective preprocessing facilitates more efficient analysis. In this paper, the authors discuss ‘prodata’, a Python library specifically developed and designed to automate and simplify common data preprocessing tasks. ‘prodata’ includes features for missing data imputation, outlier treatment, categorical data encoding, and data visualization. The results demonstrate that the execution time for the ‘prodata’ functions averages just 0.9133 seconds, significantly less than the time required for manual coding. The ‘prodata’ library remarkably reduces the average number of code lines needed for each preprocessing step to just one as compared to manual preprocessing which requires at least six lines of code per step. By offering a user-friendly interface, ‘prodata’ abstracts the complexities of data preprocessing, making it accessible to both novice and experienced data scientists alike. This work highlights the potential of ‘prodata’ to enhance the efficiency of data preprocessing.

Read the paper · More papers on PaperTik