The ML Data Prep Zoo
Vraj Shah, Arun Kumar · 2019
Data preparation (prep) time is a major bottleneck for many ML applications. It is often painful grunt work that is handled manually by data scientists, reducing their productivity and raising costs. It is also a roadblock for emerging AutoML platforms. We envision a new line of community-driven research to tackle this bottleneck based on a simple philosophy: use ML to semi-automate data prep for ML. For impactful research on this problem, we believe the major impediment is not new algorithms or theory but rather common task definitions and benchmark labeled datasets. To this end, we formalize a few major data prep tasks for ML over structured data as applied ML tasks. We discuss research challenges in scaling up data labeling, defining accuracy metrics, and creating practical tool support. We present a case study of our progress on a key data prep task: ML schema inference. Finally, we propose a public "zoo" of labeled datasets and pre-trained ML models for data prep tasks to act as a community-led repository for further research on this problem.