Unlocking the AMD Neural Processing Unit for ML Training on the Client Using Bare-Metal-Programming Tools
A. Rosti, Michael Franz · 2025
We present an optimized implementation of GPT-2 training (fine-tuning) that harnesses AMD's Neural Processing Unit (NPU) for improved throughput and power-efficiency. We use a low-level programming framework, enabling a close mapping of our application to the hardware.11Thank you to Joseph Melber, Kristof Denolf, and Phil James-Roxby from Advanced Micro Devices Inc., for their guidance, help and contributions.