A task offloading and batch scheduling framework for Edge-assisted inference
Dimitrios Spatharakis, Christos Pelekis, Dimitrios Dechouniotis, Symeon Papavasileiou · Future Generation Computer Systems · 2025
Real-time processing of inference tasks generated by resource-constrained devices in an Edge Computing environment demands carefully designed solutions that guarantee the performance and achieve a high level of accuracy. To reduce transmission time, inference tasks are often compressed to reduce the transmission time and offloaded to the edge infrastructure for parallel batch processing in GPUs. In this setting, an interesting tradeoff arises, characteristic of Approximate Computing, where the quality of inference and the system’s end-to-end latency are competing objectives. In this paper, we formulate a joint optimization problem to maximize the quality of inference while minimizing the overall latency for the GPU-enabled batch processing of inference applications. The optimization problem is NP-hard, and we split it into two subproblems to obtain optimal values for the compression of offloaded tasks and select the offloading strategy that minimizes the total latency. By carefully examining the results of the compression problem, we identify that compressing the tasks in such a way to arrive simultaneously for remote processing significantly increases the performance of batch processing. To compute an offloading strategy, we employ a semidefinite relaxation (SDR)-based approach and a randomized mapping to obtain feasible solutions. Therefore, we design an iterative alternating algorithm to solve both problems and obtain a near-optimal solution in polynomial complexity. Simulation results indicate that the proposed framework outperforms all compared solutions by reducing the total cost by 50%.