Fireworks: Reproducible Machine Learning and Preprocessing with PyTorch
Saad M. Khan, Libusha Kelly · The Journal of Open Source Software · 2019
Here, we present a batch-processing library for constructing machine learning pipelines using PyTorch and dataframes.It is meant to provide an easy method to stream data from a dataset into a machine learning model while performing reprocessing steps such as randomization, train/test split, batch normalization, etc. along the way.Fireworks offers more flexibility and structure for constructing input pipelines than the built-in dataset modules in PyTorch (Paszke et al., 2017), but is also meant to be easier to use than frameworks such as Apache Spark (Zaharia et al., 2016).