Scaling LLM Inference Architectures: A Performance Analysis for Chatbot Applications
Aditi M Jain, Ayush Jain · 2025
This paper presents a systematic evaluation of architectural patterns for Large Language Model (LLM) inference in production chatbot applications, addressing the critical challenge of balancing performance, cost, and scalability. We analyze five distinct architectures-monolithic, microservices, edge computing, event-driven, and hybrid edge-microservices-using a novel experimental framework with 100 concurrent users as a baseline. Our methodology incorporates precise measurements of latency profiles, throughput, resource utilization, and cost metrics, employing GPT-3.5-turbo with vLLM optimization. Key findings reveal that hybrid edge-microservices architecture offers 46% lower P99 latency and 67% higher throughput compared to monolithic approaches, while edge computing demonstrates 37% lower CPU usage. We introduce a scaling factor analysis methodology for accurate performance predictions at larger scales, validated through controlled experiments. This research contributes: (1) a systematic evaluation methodology for LLM inference architectures, (2) empirical evidence for architectural decision-making, (3) novel scaling factors for performance prediction, and (4) a detailed cost-benefit analysis across architectural patterns. These insights advance the scientific understanding of LLM deployment strategies and provide crucial guidance for both researchers and practitioners in optimizing large-scale neural inference systems.