Gen-AI in a Bottle: Experiments with LLMs to Generate HPC Kernels

Upasana Sridhar, Elliott Binder, Tze Meng Low · 2025

Numerical libraries derive performance from highly specialized code – known as kernels/microkernels – written by experts. Reliance on a small group of experts poses challenges to the portability and adaptability of high-performance code. This work explores the use of Large Language Models (LLMs) trained on code as a medium to automatically generate high-performance code. We use Code-LLaMA as a representative model and detail its strengths and weaknesses as it pertains to writing high-performance kernels. Then, we use these insights to describe prompting strategies to use code-LLMs to generate high-performance kernels. We use the example of matrix-multiplication kernels and present the portability of the generated code across different datatypes as well as different ISAs. We show that by encoding expertise into prompts, LLMs- alone can be used to generate kernels that achieve 99% of the peak throughput of the target hardware. By combining LLMs with a python-based code generator, we achieve more than 90% of peak throughput automatically, on par with expert-written kernels. This preliminary finding suggests that, with more specialization, LLMs can be used to automatically generate HPC code.

Read the paper · More papers on PaperTik