Language Models at the Edge: A Survey on Techniques, Challenges, and Applications
Siem Hadish, Velibor Bojković, Moayad Aloqaily, Mohsen Guizani · 2024
The invention of the transformer architectures has spurred Large Language Models (LLMs) to the forefront of artificial intelligence, driving state-of-the-art advancements across various domains. However, the huge scale of these models, often comprising hundreds of billions of parameters, presents significant challenges in terms of computational resources for both training and deployment. Consequently, the development and implementation of LLMs have largely been confined to major technology corporations. Recent research has focused on addressing these limitations by developing compact LLMs capable of operating on edge devices, rather than relying on centralized server infrastructure. This approach offers numerous benefits, including reduced computational costs, lower latency, enhanced security, and the potential for task-specific, specialized models. This paper examines the latest research trends and developments in the field of compact LLMs, exploring their applications and discussing the challenges associated with their implementation. Our investigation aims to provide insights into the future landscape of accessible and efficient language models.