RecompGPT: Generative Pre-Trained Transformers-Assisted Interactive Human Gaze Pattern Learning and Distribution Modeling for Scene Recomposition
Shang Wang, Nassiriah Binti Shaari, Nur Sauri Yahaya, Liu Hao · IEEE Access · 2025
We introduce a cutting-edge, GPTa-assisted approach to optimizing visual scene retargeting by learning human gaze behaviors. RecompGPT leverages the power of Generative GPT to model human gaze dynamics, providing an advanced mechanism for intelligent image reconstruction. The framework incorporates a hierarchical structure, utilizing the BING objectness metricbto identify key visual elements in exhibits. Through the integration of multimodal data, RecompGPT employs a novel Locality-Preserved and Observer-Like Active Learning (LOAL) strategy to incrementally generate Gaze Shift Paths (GSPs). LOAL is an active-learning algorithm that selects multiple representative patches from each scene image. It simultaneously preserves the local distribution of image patches and chooses representative (or visually salient) regions in a way that mimics human gaze allocation. Then, we deploy GPT to learn the distribution of the initial human gaze fixation toward different sceneries. Afterward, these GSPs are refined via a multi-layer aggregation algorithm that encodes deep feature representations into a Gaussian Mixture Model (GMM) to model the distribution of human gaze patterns. The learnedGMMguides the scene recomposition process by maximizing the posterior estimation. Empirical evaluations, including user studies, demonstrate RecompGPT’s superiority over existing methods, achieving a 4.39% improvement in precision and reducing testing time by 50%. RecompGPT harmonizes cutting-edge algorithmic efficiency with human-centered aesthetics, pushing the boundaries of AI-driven scene analysis and visual recomposition. It enhances the interactivity and immersion of virtual reality, offering a next-generation approach to scene recomposition that aligns with human gaze patterns and cognitive preferences.