AttenGaze: an attention-guided gaze segmentation model for efficient human-computer interaction

Hongming Shao, Juan Li, Xiaoyi Wang · 2025

Real-time gaze-guided image segmentation is vital for HCI, yet models like SAM are too costly for edge deployment. Existing methods neglect user attention’s localized nature, failing to balance generalization and real-time performance. We propose AttenGaze, an attention-guided model with a two-stage framework. First, a Region-Aware Input Reduction module fuses gaze, temporal, and image data to predict ROI, cutting 74% redundant pixels on GAZE-VIPSeg datasets. Then, an Efficient Low-Resolution Encoder processes the ROI, reducing complexity while maintaining accuracy. Experiments show 96% mAP in ROI selection, 60ms latency on RK3588, 4× faster than MobileSAM, and 11% higher mIoU than EfficientSAM. This work resolves the open-world generalization-real-time dilemma, enabling efficient HCI on resource-constrained devices.

Read the paper · More papers on PaperTik