Memory Efficient with Parameter Efficient Fine-Tuning for Code Generation Using Quantization
Purnawansyah Purnawansyah, Zahrizhal Ali, Herdianti Darwis, Lutfi Budi Ilmawan, Sitti Rahmah Jabir, Abdul Rachman Manga · 2024
Code Large Language Models (Code LLMs) such as Code LLaMa and StarCoder have exhibited outstanding proficiency in tasks required for specific tasks like code generation. Several conducted research to similar task by utilizing fine-tuning techniques from state-of-the-art base models for more specific related task. However, due to the cost limitations and limited computing resources, performing fine-tuning from large language models is excessively high. In this study, we utilized Low-Rank Adaptation (LoRA) for base large language models such as LLaMA-2 and Phi-1.5, which uses trainable rank decomposition matrices. Furthermore, we injected Quantized LoRA (QLoRA) to help reduce memory usage while training the model and analyzed the contribution to GPU usage. Notably, our findings reveal that employing these techniques for fine-tuning on small datasets yields cost-effective and viable alternatives for language-related tasks, showcasing competitive performance compared to state-of-the-art models like CodeLLaMa 7B substantiated by lower train loss achieved in our experiments.