LT-StyleCap: Lightweight Stylized Image Captioning With Dynamic Token Pruning and Factorized Style Control

Abhinav Kumar, Gangothri Sanil, Krishna Prakash · IEEE Access · 2026

Modern Multimodal Large Language Models generate impressive image captions, but their massive parameter counts and high inference latencies make them difficult to deploy on edge devices. Traditional lightweight models solve the latency problem but generate stylistically uniform and, factually constrained descriptions. LT-StyleCap bridges this gap by proposing a 75.31 M-parameter architecture that generates factual, humorous, and romantic captions in a single forward pass, achieving end-to-end inference at 43.17 ms per image on a single GPU. A Stabilized Dynamic Token Pruner (DTPStabilized) guarantees a semantic floor of visual tokens, reducing encoder FLOPs by 35% without inducing subject hallucinations. Text generation is anchored to the visual scene through an Efficient Semantic Fusion (ESF) module that integrates soft attribute tags into the spatial feature stream via multi-head self-attention. For stylistic steering, a Factorized Style Control (FSC) layer replaces the standard output projection with a three-factor matrix ( $\mathbf {U}\mathbf {S}_{\ell }\mathbf {V}$ ), whose per-style interaction core is conditioned at the inference time without requiring parameter swaps. Training combines Style Contrastive Learning with 18,000 hard synthetic captions distilled from BLIP-2, followed by a Stratified Balanced Supervised Fine-Tuning (SFT) stage on a 9,588-token vocabulary. Evaluated on the Flickr8k + FlickrStyle10k benchmark, LT-StyleCap achieves Style-Accuracy of 100.0%, 53.3%, and 62.7% for factual, humorous, and romantic tracks, respectively, while maintaining strong lexical diversity (Distinct-2 up to 64.5%). Zero-shot style prompting applied to a 3.7 B-parameter BLIP-2-OPT-2.7B model fails in this task, achieving only 3.9% humor Style-Accuracy. LT-StyleCap outperformed the $\sim 50\times $ larger model by 49.4 percentage points while operating at $65\times $ lower latency.

Read the paper · More papers on PaperTik