Deep Active Learning for Address Parsing Tasks with BERT

Berkay Güler, Betül Aygün, Aydın Gerek, Alaeddin Gürel · 2023

Deep learning models tend to perform better with larger datasets. With decreasing data handling costs, researchers have the means to gather and store vast amounts of unlabeled data. Supervised learning, on the other hand, requires training data to be labeled by annotators. However, high annotation costs pose challenges to labeling an optimum portion of the available data. One proposed method to mitigate this problem is to employ active learning (AL). AL strategies use a machine learning model to select the most informative and representative samples among unlabeled data points. Here, we demonstrate the effectiveness of uncertainty-based active learning strategies, including a new strategy, for address parsing with a BERT model on an in-house Arabic address dataset manually annotated for two different tasks. We compare AL methods with random sampling and longest-sentence baselines. We show that AL strategies' usefulness greatly depends on dataset characteristics, being less effective on datasets with fewer classes. We conclude that AL for address parsing with BERT decreases annotation costs, if measured in the number of queries. Yet, due to AL methods' tendency to select longer queries, some strategies may increase labeling costs, measured in the total number of words.

Read the paper · More papers on PaperTik