Self-Supervised Transformer Networks for Error Classification of Tightening Traces
Dennis Wilkman, Lifei Tang, Kateryna Morozovska, Federica Bragone · 2022
Transformers have shown remarkable results in the domains of Natural Language Processing and Computer Vision. This naturally raises the question of whether the success could be replicated in other domains. However, due to Transformers being inherently data-hungry and sensitive to weight initialization, applying the Transformer to new domains is quite a challenging task. Previously, the data demands have been met using large-scale supervised or self-supervised pre-training on a similar task before supervised fine-tuning on a target downstream task. We show that Transformers are applicable for the task of multi-label error classification of trace data and that masked data modelling based on self-supervised learning methods can be used to leverage unlabelled data to increase performance compared to a baseline supervised learning approach.