SwiftDP: An Efficient Framework for Automated Data Preparation Pipeline Generation
Liangwei Li, Yiyi Zhang, Ning Wang · 2025
As a crucial step of machine learning, data preparation is the most time and energy consuming task for data scientists, entailing several data processing techniques to improve the performance of output results for ML models. However, end-to-end AutoML researchers primarily focus on automated ML pipelines consisting of algorithm selection and hyperparameter tuning, while automated data preparation has not been widely explored and applied. In this paper, we propose SwiftDP, an efficient framework for automated data preparation based on Monte Carlo Tree Search. To guide the search more effectively, a fully connected neural network with an attention mechanism is designed to estimate the subsequent maximum performance gain of each tree node. In addition, in order to reduce search space and improve system efficiency, two optimization strategies, meta-learning and accelerated training strategy, are used to determine the type and order of tasks in the data preparation process in advance, and speed up the pipeline creation process. We have built SwiftDP as an open-source Python library with intuitive APIs and demonstrated its better performance in 10 seconds than SOTA systems in 1 hour.