Interest Region Prediction in XR Dramas Using Transformer Networks
Siqi Lei, Ronghua Ke, Jiamin Qi · 2025
In this paper, we propose a Transformer-based model for predicting user interest regions in XR short drama scenarios by integrating visual features and simulated gaze data. Using the MouseGaze dataset as a proxy for eye-tracking, we construct time-aligned sequences of keyframes and attention maps, and encode them through a spatio-temporal Transformer framework. Our model captures the dynamic evolution of user focus via position-aware self-attention, and outputs both saliency maps and bounding boxes for fine-grained interest prediction. Experimental results demonstrate that our method significantly outperforms existing baselines-including CNN+LSTM and ViT-in terms of AUC, NSS, and CC metrics. Visual comparisons further confirm its ability to predict attention regions that closely match human behavior. This study presents a practical and reproducible approach for visual attention modeling in XR environments, offering promising applications in gaze-guided storytelling, adaptive scene control, and immersive interaction design.