Revisiting Text-Guided Style Transfer Through Mamba: A Fast and Memory-Efficient Architecture

Filippo Botti, Teresa Calzetti, Alex Ergasti, Tomaso Fontanini, Andrea Prati · IEEE Access · 2026

Text-guided style transfer aims to stylize a content image according to a textual description while preserving the image structure. Unfortunately, many recent approaches rely on either attention-based fusion or diffusion backbones. On one hand, transformers require high memory consumption due to their quadratic complexity, while diffusion-based methods require iterative denoising, making inference slow and computationally expensive. In this work, we propose a lightweight architecture based on Mamba state-space models. Our key idea is to perform cross-modal fusion inside the Mamba dynamics by conditioning the state-space parameters on the text embedding while processing content tokens, avoiding attention and enabling linear-complexity fusion. Additionally, we represent the content image as a sequence of multi-scale frozen VGG features instead of learned patch embeddings, reducing sequence length and improving content encoding. Experiments at 512×512 resolution show that our method achieves strong prompt alignment while maintaining high structural similarity to the input and it provides a favorable quality-efficiency trade-off compared to open baselines. Code is available at https://github.com/FilippoBotti/TGSTM.

Read the paper · More papers on PaperTik