TASCA : Tool for Automatic SCalable Acceleration of ML pipelines✱

Archisman Bhowmick, Mayank Mishra, Rekha Singhal · 2024

Data Scientists use Python for building ML pipelines including pre-processing on data for cleansing and transformation. The computational overheads, due to performance anti-patterns, on data processing can be expensive, especially on large data size. FASCA [10] is a framework to identify such performance anti-patterns, with their corresponding performant versions, from ML pipelines. However, it needs human intervention to generate the performant code. Recent growth in maturity of Large Language Models (LLM) for code generation motivated us to exploit them for automating the process of transforming performance anti-patterns with their performant versions, the feasibility of which has been discussed in [5].

Read the paper · More papers on PaperTik