Characterizing Communication in Distributed Parameter-Efficient Fine-Tuning for Large Language Models
Nawras Alnaasan, Horng-Ruey Huang, Aamir Shafi, Hari Subramoni, Dhabaleswar K. DK Panda · 2024
Parameter-efficient Fine-tuning (PEFT) methods have emerged as powerful techniques for adapting pre-trained Large Language Models (LLMs) to specific tasks with reduced computational and memory overhead. However, despite their promising potential, there remains a gap in understanding how these methods perform in distributed computing settings. In this paper, we present a comprehensive characterization of the communication dynamics in distributed PEFT for LLMs. Our study emphasizes the crucial role of communication efficiency in the performance and scalability of PEFT methods when utilized across GPU clusters. Through systematic analysis of various PEFT techniques and LLM sizes, we have illustrated how communication overhead can significantly impact throughput, training time, and overall model performance. Our findings indicate that PEFT methods inherently reduce communication and computational burdens compared to full fine-tuning. These improvements yield up to 1.75x speedup by using PEFT methods such as Low-Rank Adaptation (LoRA) to fine-tune GPT-like billion parameter generative LLMs. We conduct this characterization study on modern GPU clusters with InfiniBand interconnect with up to 32 NVIDIA A100 GPUs. To the best of our knowledge, this is the first effort that evaluates the performance of PEFT methods in distributed computing environments. Ultimately, this work strives to provide valuable insights into the efficacy and practical implications of employing PEFT methods and facilitate the development of scalable approaches for fine-tuning large-scale models.