Large Language Models in AWS

Ravi Kishore Kodali, Yatendra Prasad Upreti, Lakshmi Boppana · 2024

Large Language Models (LLMs) have become a transformative milestone in natural language processing. The deployment of LLMs on cloud platforms like Amazon Web Services (AWS) plays a pivotal role in making their capabilities accessible to a wider audience, contributing to the democratization of advanced language processing technology. This paper reviews the evolution of cloud application architectures, focusing on transferring cloud applications across various infrastructures, offering insight for cloud engineers and architects. Throughout this paper, we dive into the intricacies of deploying Llama-2 on various AWS instance types, each tailored to specific use cases and computational requirements. The analysis highlighted trade-offs between factors such as computational speed, cost efficiency, and resource utilization, all of which play a crucial role when working with the LLM model. Additionally, we compared instance performance, emphasizing the role of max-batch-prefill tokens in enhancing response generation efficiency. The results provide a clear understanding of the impact of instance-type selection on the model’s performance, allowing informed decisions to be made based on specific language processing requirements.

Read the paper · More papers on PaperTik