Proactive Elastic Scheduling for Serverless Ensemble Inference Services
Shiyu He, Binbin Feng, Zhijun Ding · 2024
Recently, AI inference services have adopted ensemble architectures, which are widely recognized and used for their advanced performance. However, the existing ensemble inference services are mainly created and managed using the platform-as-a-service model with the static ensemble service architecture and rigid persistent resource allocation. This makes the highly heterogeneous inference requests rely on a fixed combination of basic learners and manual homogeneous resource management, resulting in insufficient precision, waste of resources, and high management costs. Serverless computing, represented by Functions-as-a-Service (FaaS), realizes transparent on-demand resource allocation to developers, which is suitable for ensemble inference services. Therefore, we propose a serverless proactive elastic scheduling solution PESEI for ensemble inference services. First, a two-level hierarchical dynamic ensemble service framework with joint model precision and overhead sensing is proposed to ensure inference precision while improving cost efficiency; second, a proactive elastic resource allocation algorithm with dynamic sensing of workload patterns is proposed to optimize the quality of service and cost efficiency for heterogeneous base learners; based on this, a prototype system is designed and developed to support the autonomous management of ensemble inference services, which implements the benign combination and adaptation of dynamic service architecture and elastic resource management. Real cluster experiments on public datasets demonstrate the effectiveness and robustness of PESEI, providing a new solution for building serverless ensemble inference services.