A Study on Named Entity Recognition with Different Word Embeddings on GMB Dataset using Deep Learning Pipelines

Surya Teja Chavali, Charan Tej Kandavalli, T M Sugash, Deepa Gupta · 2022 13th International Conference on Computing Communication and Networking Technologies (ICCCNT) · 2022

Many natural language applications, such as QA tools, information retrieval, text summarization, and machine translation, are built on top of Named Entity Recognition (NER). The primary task of Named Entity Recognition is to classify all objects (Named Entities) in a given text sample into specified classes such as person, location, and organization names. This work aims to create a NER model for the popular Groningen Meaning Bank (GMB) dataset, comprising tens of thousands of unprocessed and tokenized texts. The dataset has been published by the University of Groningen and updated by developers at IBM in the year 2020. The proposed system is designed into 7 different models – passing the embedding vectors created by keras embedding (Python), BERT, Glove, Word2Vec, fastText, and other two involving Character Embedding along with the Attention layer into a BiLSTM-CRF network. These techniques were employed to compare and analyze the model performances against each other for a better understanding.

Read the paper · More papers on PaperTik