Efficient Token Augmented Fine-Tuning
Lukas Jonathan Weber, Krishnan Jothi Ramalingam, Matthias Beyer, Alice Kirchheim, Axel P. Zimmermann · 2023
Continual pretraining related performance gains come at significant costs in terms of money, computational time, and the environment. In this paper, we present the method of Efficient Token Augmented Fine-Tuning (TAFT) which is an affordable and fast alternative for resource-intensive and time-consuming continual pretraining activities. The vocabulary of the tokenizer has been augmented with task-specific tokens and the token embedding has been initialized using an intelligent strategy. The model size of RoBERTa has increased by less than 1 % due to vocabulary augmentation, and an increase in performance in downstream tasks by up to 2 % fl-score has been generated. We achieved new state-of-the-art results for the datasets Hyperpartisan and HLGD. The task-specific token identification and embedding initialization steps have taken a maximum of 15 minutes with a 448 GB memory 24 core CPU for all the datasets experimented with in this work.