Single or Double: On Number of Classifier Layers for Small Language Models
Muhammed Cihat Ünal, Mounes Zaval, Aydın Gerek · 2023
Pretrained transformer language models such as BERT have become very popular among NLP practitioners, and as such have been applied to a variety of tasks. BERT's strength derives from producing contextual vector representations for text, which are then traditionally used with a linear classifier (a single fully connected layer). One of the less famous NLP tasks that BERT has been applied to is address parsing, which deals with breaking an address into its various components, such as street name or door number. In an earlier study on address parsing it was observed that replacing the traditional single layer linear classifier with a double layer MLP (Multi-Layer Perceptron) could yield an increase in performance metrics if the language model providing the underlying representation had low capacity (as assumed by its low number of parameters). In this study, we examine whether this observation carries to the much better known NLP task of text classification.