Scaling Address Parsing Sequence Models through Active Learning
Helen B. Craig, Dragomir Yankov, Renzhong Wang, Pavel Berkhin, Wei Wu · 2019
Address parsing is a critical step for map search engines. This component annotates the terms of an address query, e.g. house numbers, road names, administrative units etc., so that the address search engine can resolve the expected result. Deep recurrent models achieve state of the art performance for address parsing; however, scaling such models is problematic. They require a significant amount of term-annotated data which is expensive to acquire. In this paper, active learning significantly reduces the amount of labeled data required to train accurate address parsing models. We demonstrate the efficiency of our approach when cold-starting with human-labeled as well as synthetically-generated data.