DiLLeMa: An extensible and scalable framework for distributed large language models (LLMs) inference on multi-GPU clusters

Robby Ulung Pambudi, Ary Mazharuddin Shiddiqi, Royyana Muslim Ijtihadie, Muhammad Nabil Akhtar Raya Amoriza, Hardy Tee, Fadhl Akmal Madany, Rizky Januar Akbar, Dini Adni Navastara · SoftwareX · 2026

The increasing demand for scalable and responsive Large Language Model (LLM) applications has accelerated the need for distributed inference systems capable of handling high concurrency and heterogeneous GPU resources. This paper introduces DiLLeMa, an extensible framework for distributed LLM deployment on multi-GPU clusters, designed to improve inference efficiency through workload parallelization and adaptive resource management. Built upon the Ray distributed computing framework, DiLLeMa orchestrates LLM inference across multiple nodes while maintaining balanced GPU utilization and low-latency response. The system integrates a FastAPI -based backend for coordination and API management, a React -based frontend for interactive access, and a vLLM inference engine optimized for high-throughput execution. Complementary modules for data preprocessing, semantic embedding, and vector-based retrieval further enhance contextual relevance during response generation. Illustrative examples demonstrate that DiLLeMa effectively reduces inference latency and scales efficiently.

Read the paper · More papers on PaperTik