Performance modelling and analysis of IaaS CCP architecture
Kabiru M. Maiyama, Demetres D. Kouvatsos · Expert Systems with Applications · 2026
Queueing Network Models (QNMs) constitute robust quantitative tools for evaluating and predicting the Quality of Service (QoS) of virtual machine (VM) provisioning in Cloud Computing Platforms (CCPs). An extended application-driven open QNM at equilibrium with random routing and a First-Come-First-Served (FCFS) rule is proposed for the performance modelling and analysis of Infrastructure as a Service (IaaS) CCP architecture. It is assumed that the QNM has a bursty external arrival process of clients requesting VMs according to a Compound Poisson Process (CPP) with geometrically distributed batches, or equivalently, highly variable Generalised Exponential (GE) interarrival times. Moreover, it consists of L ( L ≥ 1 ) queueing stations with infinite buffer capacity and c i ( c i ≥ 1, i = 1, 2, …, L ) GE-type servers. By means of a generic maximum entropy (ME) product form approximation, the proposed QNM is decomposed into individual GE/GE/c i , i = 1, 2, …, L queueing stations with GE-type interarrival and service times, each of which can be analysed in isolation. Consequently, closed-form expressions for key performance metrics of each queueing station i, i = 1, 2, …, L of the GE-type QNM are devised, such as those of throughput, server (resource) utilisation, mean response time and steady state probability of number of VM requests by clients. Moreover, typical numerical experiments are carried out, based on exponential memoryless (M−type), a family of two-phase Hyperexponential-2 (H 2 ), and GE interarrival and service time distributions of the QNM, using MATLAB and JMT tools as appropriate, which shows GE-type yield pessimistic bound or worst-case performance. The evaluation of the performance metrics of the QNM and the prediction of related bottlenecks were presented. This will provide vital insights for capacity planning at the design and development stage of new IaaS CCP architectures, as well as for tuning and upgrading existing ones to facilitate effective processing and leasing of VMs to requesting clients. We integrate our QNM into an expert system workflow comprising analytical outputs (such as overall delay W̄, station(s) utilisations ρ i , bottleneck identity) feed simple, auditable rules that recommend scale out, throttling, or routing adjustments. Sensitivity results are translated into proactive policies with clear thresholds.