EM-SAM: Eye-Movement-Guided Segment Anything Model for Object Detection and Recognition in Complex Scenes

Jinqi Li, Yang Yu, Junfan Zhou, Chinan Wang, Ling‐Li Zeng · 2024

Over the past few decades, object detection and recognition systems have made great strides relying on convolutional neural networks (CNNs). However, these methods perform poorly in small object detection and complicated natural scenes. There are still plenty of opportunities for improvement in enhancing small object features extraction and eliminating the effects of complex scenes. Compared with computer, human can automatically ignore redundant information in complex scenes, focusing attention on suspected objects. To this end, we proposed an Eye-Movement-Guided Segment Anything Model for Object Detection and Recognition in Complex Scenes Specifically, the framework includes eye movement state classification for acquiring gaze points, object detection based on gaze point segmentation to remove the influence of complex scenes, and target recognition for lowering the confidence threshold. In the object detection using Segment Anything Model (SAM), the object detection success rate was 97.6% ± 3.19% when the false detection rate was 9.62% ± 6.45%. Compared to the baseline, the Recall under the object recognition increased from 85.71% to 91.83%, and the MAP increased by 5.33%. Meanwhile, The average time for selection of the framework's targets was 1.35s, improving the user's ability to interact with the environment.

Read the paper · More papers on PaperTik