Enhancing Large Language Models for Telecom Networks Using Retrieval-Augmented Generation

Nasik Sami Khan, Md Mahibul Hasan, Md. Shamim Towhid, Saroj Basnet, Nashid Shahriar · 2024

This paper presents a comprehensive approach for fine-tuning large language models (LLMs) for domain-specific tasks in the telecommunications field. We utilize a dataset with 1,827 multiple-choice questions (MCQs) from 3GPP standard documents. A publicly available LLM named "Phi-2" is used to answer the MCQs correctly. We develop a Retrieval-Augmented Generation (RAG) pipeline to improve Phi-2 model's performance. The RAG pipeline comprises document segmentation, synthetic question-answer (QA) generation, custom fine-tuning of the embedding model, and incremental fine-tuning of Phi-2. Our experiments show that accuracy greatly increased by combining all the above-mentioned steps in the RAG pipeline. The proposed approach outperforms the baseline Phi-2 model by 45.20% in terms of accuracy. This study identifies the limitations of instruction fine-tuning in specialized fields and explores the possibility of using sophisticated data processing with fine-tuned models to improve performance even more.

Read the paper · More papers on PaperTik