Semantic Representations in Text Data
Triveni Lal Pa, Madhu Kumari, Tajinder Singh, Mohammad Ahsan · International Journal of Grid and Distributed Computing · 2018
Automatic text mining processes and other sophisticated natural language processing constructs need realistic representations of text/documents which embed semantics efficiently.All the representations work on the notion that every data contains different explanatory factors (attributes).In this article, we exploit these explanatory factors to study and compare various semantic representation methods for text documents.The article critically reviews recent trends in the area of semi-supervised semantic representations, covering cutting-edge methods in distributed representations such as embeddings.This article gives a broad and synthesized description of various forms of text representations, presented in their chronological order ranging from BoW models to the most recent embeddings learning.Conclusively, various findings taken together provide valuable pointers for researchers looking to work in the field of semantic representations.In addition, the article also shows that one need to develop a model for learning universal embeddings in unsupervised/semi-supervised settings that incorporate contextual as well as word-order information, with language independent features and which would be feasible for large dataset.