Active Learning for Turkish Text Classification
Ali Osman Berk Şapcı, Öznur Taştan, Reyyan Yeniterzi · 2020
Many natural language processing (NLP) applications are solved in a supervised learning framework, which often requires a sufficiently large amount of labeled data. However, the labeled data is scarce as the process of obtaining class labels is often costly and/or time-consuming and sometimes requires indepth expert knowledge. Thus, obtaining labeled sets is a bottleneck. On the other hand, unlabeled data is abundant and easy to access in a variety of domains. The active learning paradigm addresses this challenge by effectively selecting informative and/or representative instances from the unlabeled pool to be labeled. In this way, active learners aim to reduce the cost of label acquisition without sacrificing model performance. In this work, we apply conventional active learning techniques on various Turkish text classification tasks. Experiments demonstrate that active learning helps to attain good performances while reducing the required labeled data.