Gaze Direction Classification Using Vision Transformer
Shogo Matsuno, Daiki Niikura, Kiyohiko Abe · 2023
In this paper, as a study of basic technology for developing input interfaces using eye movement, we proposed a method for constructing a model for estimating eye direction using the Vision Transformer and evaluated its performance. Appearance-based gaze estimation methods can provide relatively robust and accurate gaze estimation even for video images captured by ordinary video cameras, but they require a certain amount of computational resources to process with sufficient accuracy. Therefore, in this paper, we aim to develop a machine learning model that detects only eye movement, not the gazing point, to simplify inference for the practical use of calibration-free and low-cost eye input interfaces. The proposed method constructs a gaze direction estimation model by fine-tuning a large-scale pretrained model of Vision Transformer using a constructed dataset. The dataset is constructed by extracting the face region from a frontal image of a person captured by a common webcam as a still image for each frame, and then cropping the region near the eyeballs. In addition, to evaluate the performance of the gaze direction estimation model constructed by the proposed method, we conducted an experiment to compare it with a classification model constructed by a conventional CNN - based method. The experimental results showed that the proposed ViT model has a higher classification performance than the conventional CNN model by approximately 6.0 points in terms of Accuracy and 5.9 points in terms of Macro Average F -value, confirming the overall improvement of the classification performance. This indicates that the gaze direction estimation model constructed using the proposed method is effective as a fundamental technoloav for gaze input interfaces.