Exploring Cooperative Caching for AI-Generated Content Inference in Edge Networks

Meng Tian, Xingyi Cai, Yunfeng Zhao, Chao Qiu, Xiaofei Wang · 2025

Artificial Intelligence Generated Content (AIGC) services based on text-to-image generation tasks have driven change in the AI industry in recent years. Diffusion models are widely used in AIGC services for generating high-quality images from complex prompts. However, the process of generating diffusion models requires a large number of autoregressive denoising steps, posing significant challenges related to service latency and privacy issues, especially in resource-constrained edge devices. To further improve the quality of service (QoS) of edge AIGC services, we design a multi-edge cooperative caching mechanism based on the idea of reusing early intermediate results of similar prompts to reduce denoising steps. First, a collaborative filtering algorithm based on prompt similarity is developed to analyze user preferences. The designed multi-agent reinforcement learning-based cache management (MARLCM) algorithm utilizes the preference data as input and determines the caches that need to be replaced in the cache space. Secondly, we propose an adaptive edge server cooperation strategy that instructs servers to build efficient cache pools based on server similarity and workload differences. The experiment conducted on a high-resolution generation task with multiple diffusion cache frameworks shows that the proposed method reduces task latency by 23.1 %, and achieves up to 1.57 times higher cache hit rate compared to existing methods.

Read the paper · More papers on PaperTik