Exploring Shared Large Language Models: Early Insights into Scalability and Efficiency in AI Assistant and Agent Deployment

Arvid Kok, Antonio Carvalho, Michael D. Street · 2025

The deployment of Large Language Models (LLMs) is rapidly expanding across diverse applications, necessitating cost-effective and resource-efficient strategies to optimize their usage. This paper investigates the scalability, efficiency, and performance trade-offs of sharing LLMs across multiple applications, addressing critical challenges such as GPU limitations, concurrency management, and latency optimization. Using three experimental setups ranging from consumer-grade GPUs to high-performance cloud infrastructure, we examine the interplay between prompt size, model size, and concurrency on metrics like latency, throughput, and GPU utilization. Our findings reveal that shared LLM architectures significantly enhance resource efficiency, with concurrency improving throughput by 2x to 4x for longer prompts and over 20x for shorter batched prompts. However, memory constraints impose limitations on scalability, particularly for large models and extended prompts, where latency increases linearly with context length. Practical recommendations include tailoring GPU configurations to balance memory and compute demands, leveraging batching for optimal utilization, and mitigating latency through caching and load balancing. This study underscores the strategic value of shared LLMs in reducing costs and enhancing scalability for multi-application scenarios, particularly in domains with constrained resources, such as defense. The results provide actionable insights into deploying shared generative AI systems efficiently while paving the way for future exploration of advanced optimization techniques. This paper was originally presented at the NATO Science and Technology Organization Symposium (ICMCIS) organized by the Information Systems Technology (IST) Panel, IST-209-RSY-the ICMCIS, held in Oeiras, Portugal, 13–14 May 2025.

Read the paper · More papers on PaperTik