Pocket: ML Serving from the Edge

Misun Park, Ketan Bhardwaj, Ada Gavrilovska · 2023

One of the major challenges in serving ML applications is the resource pressure introduced by the underlying ML frameworks. This becomes a bigger problem at resource-constrained, multi-tenant edge server locations, where it is necessary to scale to a larger number of clients with a fixed resource envelope. Naive approaches which simply minimize the resource budget allocation of each application result in performance degradation that voids the benefits expected from operating at the edge.

Read the paper · More papers on PaperTik