Evaluating three approaches to binary event-level agreement scoring. A reply to Friedman (2020)

Raimondas Zemblys, Diederick Christian Niehorster, Kenneth Holmqvist · Behavior Research Methods · 2020

Recently, Friedman (2020) published a letter in which he claims there are three errors and two problems in our paper "gazeNet: End-to-end eye-movement event detection with deep neural networks" (Zemblys et al., 2019).Here we respond to these claims by Friedman, namely that improper data were used for Zemblys et al. (2019) and that performance was improperly evaluated.Let us first recap what we presented in Zemblys et al. (2019).gazeNet is a method that takes an existing eyemovement data set that has been labeled (through handcoding or by any other means) and trains a classifier to reproduce this event coding.The goal of gazeNet, as for any machine learning-based classifier, is to produce coding similar to what it observed during training.As such, the performance of classifiers like gazeNet is evaluated on other labeled data that was not seen during training, and the classifier is said to perform well if it is able to produce high agreement with the testing set (i.e., similar coding as the testing set).As such, the classifier can be trained on any input data, regardless of its quality, since the success of a classifier is determined by its performance on the testing set.In Zemblys et al. (2019), we used the procedure we proposed and trained a specific classifier using part of the lund2013-image data set (Larsson et al., 2013, see "Data" section in Zemblys et al. (2019) for detailed description),

Read the paper · More papers on PaperTik