Cache-Aware Client-Side Request Planning for Black-Box LLM APIs

Prajjwal Chittori · Zenodo (CERN European Organization for Nuclear Research) · 2026

Most applications consume large language models through black-box, per-token-billed APIs. The developer does not own the model weights, cannot change server kernels, and is billed for every token the server processes. A common folk belief holds that this bill can be lowered by relocating work to the client, for example by tokenizing on the device and shipping token identifiers, or by prompting in a more information-dense language. We show that the relevant invariant rules these out: a consumer pays for the tokens the server processes, so moving computation to the client changes the bill only if it reduces billable tokens or calls at fixed task quality. Under that constraint the space of useful client-side interventions is small and well defined. We give a cost model for it, a taxonomy that separates the few legal moves from the many that are null or impossible, and we identify one intervention that is simultaneously lossless (no task-quality cost) and local-compute-free: aligning requests to the provider's own prompt cache. We present a greedy prefix-clustering scheduler that reorders and times a request stream to maximize provider-cache hits under a short time-to-live, and we characterize the regime in which it helps. In simulation under representative public cache parameters, the scheduler reduces billed cost by up to 60 percent on an agentic tool-use workload with no quality loss, dominating prompt compression and semantic caching on cost, local compute, and quality at the same time. We also show where it stops helping: when the request rate is so slow that even back-to-back same-prefix calls fall outside the cache window, and on human-paced chat where within-session gaps are not reorderable.

Read the paper · More papers on PaperTik