A survey on word representation in natural language processing

Mustafa Nabeel Salim, Ban Shareef Mustafa · AIP conference proceedings · 2022

This survey explores several techniques and language models for the representation of raw text data. These models are considered the feature extraction phase in Natural Language Processing (NLP). The traditional approaches such as Bag-of-Words, N-Gram, and TF-IDF are simple to implement. Moreover, the recently proposed models in the literature such as Word2Vec, Glove, Fast text, ELMO, and BERT have a different effect on semantics and syntactic for the representation of words and sentences. Also, these models are included in word embedding terminology. Finally, our survey presents a comprehensive review of the state-of-the-art in the field of NLP and we believe this is of interest for the researchers in this field.

Read the paper · More papers on PaperTik