Distributed and Controllable Mobile Text-to-Image Generation With User Preference Guarantee

Yuxin Kong, Peng Yang, Xue Qin, Jizhe Zhou, Xuemin Shen · IEEE Transactions on Mobile Computing · 2025

In this paper, we investigate controllable mobile text-to-image generation at scale, considering diverse user preferences. In particular, we observe that, by incorporating visual conditions (e.g.,Canny maps and depth maps) as supplementary inputs alongside text prompts, fine-grained and controllable image generation could be achieved. To this end, we propose a system design for distributed and controllable mobile text-to-image generation by leveraging edge computing. This system can satisfy diverse user-specified quality preferences at reduced transmission cost through effective cooperation of mobile and edge computing. In particular, the proposed system consists of aVisual Condition Engineeringmodule and aDistributed Denoising Controlmodule. Since extensive profiling reveals that different visual conditions affect both generation quality and sensitivity to image encoding parameters, the first module selects the optimal configuration of user-specific visual condition on mobile devices. Key to this module is a Pareto Frontier-based model which subtly balances user-preferred generation quality and transmission efficiency. The second module enables collaborative generation by adaptively distributing denoising tasks between mobile devices and the edge server, according to their available computing resources. At the core of this module is an efficient deep reinforcement learning algorithm designed to optimize the dynamic distribution of denoising tasks. By integrating the deep diffusion model, this algorithm achieves superior action space exploration capabilities while maintaining fast convergence and reliable execution, thereby facilitating enhanced adaptability under variable computing resource scenarios. Extensive experimental results reveal that, the designed system can achieve a reduction in transmission cost by over 90% and enhance user satisfaction by up to 18%, with consistent performance across various diffusion models under diverse resource constraints.

Read the paper · More papers on PaperTik