Transport-Aware Performance Modeling for Big Data Workloads in Network-Constrained Hybrid Cloud Systems
Attila Kanto, Attila Csaba Marosi · IEEE Access · 2026
Accurate execution-time prediction for big data workloads in hybrid cloud environments is essential for offloading decisions at submission time.However, prediction remains challenging due to the nonlinear effects of network transport between on-premises clusters and public cloud resources, particularly bandwidth constraints and interconnect propagation latency.Existing approaches are often not suitable for submission-time offloading decisions: they either assume fixed infrastructure conditions without explicitly modeling transport effects, or require target-side pilot executions before making predictions.To address this problem, we present a transport-aware performance modeling framework for execution-time prediction in network-constrained hybrid clouds.By decoupling network transport effects from application logic and leveraging historical on-premises execution profiles, the framework enables execution-time estimation without per-query pilot runs on the target environment.The framework combines a stage-level analytical model whose components are derived from network transport principles (bandwidth-limited throughput, RTT-bound storage latency, and TCP slow-start dynamics) with a dependency-aware criticalpath aggregation to derive end-to-end job runtime predictions.Its three structural coefficients are calibrated using Non-Negative Least Squares (NNLS) to preserve physical interpretability.The framework is evaluated using the TPC-DS benchmark on Apache Spark across eleven emulated hybrid network configurations, covering 2-25 Gb/s bandwidth and 0.3-25 ms round-trip time.Across these heterogeneous network conditions, the model achieves consistently high explanatory power, with R 2 > 0.95 at the stage level and R 2 > 0.98 at the job level.The gains are most pronounced under jointly constrained bandwidth and latency conditions.In the most severe configuration evaluated (2 Gb/s, 25.46 ms), the proposed framework reduces the Mean Absolute Error (MAE) of execution-time predictions by 76.8% (from 28.20 s to 6.54 s), while increasing job-level R 2 from 0.6625 to 0.9858 compared to a transport-agnostic baseline.