A Comprehensive Study on Non-sequence and Sequence modeling word vector embedding approach for clinical text Named Entity Classification
Adyasha Dash, Manjusha Pandey, Lipika Mohanty, Ayesha Mohanty, Neha Bajpayee · 2024
Human languages are hard to interpret and can not be understood by the computer. Thus, teaching a computer to understand human language is a difficult endeavor that has only lately been made possible by the application of Natural Language Processing (NLP) combined with the recent advancements and developments in fields such as Deep Learning (DL) with a manifold improvement in Recurrent Neural Network (RNNs) and the use of Word Embeddings. The language modeling and feature learning method in NLP known as “word embedding” maps vocabulary to actual number vectors leveraging products like word2vec, GloVe, and fastText.NLP can help us create more effective deep-learning models to solve language problems. Statistical analysis, machine learning methods, and deep learning all benefit from the improved word and phrase processing that natural language processing (NLP) techniques provide.NLP turned unstructured text data into more structured data that expert systems could easily modify and evaluate. Using embeddings, it is possible to handle textual input more quickly and effectively while building robust deep-learning models. Studies have been able to successfully verify the improvements obtained as a result of the application of DL techniques and models for tackling classes of problems related to Biological Named Entity Recognition (BioNER), with impressive and promising outcomes. Various ML-based NLP tasks now routinely use deep learning (DL), which removes the requirement for task-specific feature engineering based on in-depth domain expertise and facilitates the identification of salient aspects. Deep learning techniques frequently employ neural network (NN) architecture, which can automatically deduce patterns from vector data and pick the most pertinent features. Currently, the neural network most frequently used for NLP tasks is called Long Short-Term Memory (LSTM).The discovery of long-term connections between medical entities is performed using LSTM, a specific type of Recurrent Neural Network (RNN), which also increases training accuracy overall. This chapter entails a comprehensive study of the non-sequence and sequence modeling embedding approach for clinical text corpus named entity classification and a comparative analysis on previously adopted approaches. We compared the resulting precision and other metrics such as the recall as well as the F1-Score of our model to those of other currently available models for a number of gene and protein entity categories. Our proposed approach obtains an F1-Score of 77.34% in 16 epochs. We have worked on a relatively smaller gold standard corpus and word embedding that has a slight impact on the result. Overall, due to the complexity of the bio-medical named entities, our proposed architecture has greatly improved entity extraction identification and classification, although there is still potential for improvement.