Impact of Free-Viewing Search on Transformer-Based Goal-Directed Search
Gabriel Marques, Plínio Moreno · 2024
Predicting human gaze is a complex challenge with numerous applications in computer vision. Recent research has predominantly treated free-viewing and goal-directed tasks independently. This work, however, posits that insights gained from free-viewing tasks can enhance the realism and effectiveness of goal-directed search tasks. By building on the Gazeformer model, which integrates a language model for target encoding and a transformer encoder-decoder structure, this study explores the impact of free-viewing data on goal-directed tasks. It identifies optimal methods for incorporating this information. A new model is proposed, consisting of two transformer-based architectures (a free-viewing model and a goal-directed model), where free-viewing data is utilized as an auxiliary task to improve overall performance.