Named Entity Recognition on Morphologically Rich Language: Exploring the Performance of BERT with varying Training Levels
Yuksel Pelin Kilic, Duygu Dinc, Pınar Karagöz · 2020
Named Entity Recognition (NER) is an information extraction task that aims to automatically identify named entities in a given text. Named entities are special types of nouns or noun groups that refer to specific entities including as person, location, organization, date, time, money and percentage. NER also facilitates various other Natural Language Processing (NLP) related tasks such as summarization and question answering. It is a vastly studied problem especially on English texts, however number of NER studies on Turkish is very limited. Being a morphologically rich language, Turkish has an agglutinative structure and hence automated analysis and information extraction performance is generally lower than those on English. The previous studies mostly use conventional supervised learning and sequence tagging methods, such as Conditional Random Fields (CRF). Only few studies use deep neural models for NER problem on Turkish texts. In this work, we particularly focus on the recent neural model, Google BERT, and analyze its performance on Turkish texts. In addition to fully trained BERT model, we investigate the performance of different training levels from fully trained to fully pre-trained.