Text Representation Models based on the Spatial Distributional Properties of Word Embeddings
Narendra Babu Unnam, Krishna Reddy Polepalli, Amit Pandey, Naresh Manwani · 2024
In the current digital era, about 80% of the digital data which is being generated is unstructured and unlabeled natural language text. In the development cycle of information retrieval and text mining applications, text representation is the most fundamental and critical step, as its effectiveness directly impacts the application’s performance. The existing traditional text representation frameworks are mostly frequency distribution-based. In this work, we explored the spatial distribution of word embeddings and proposed two text representational models. The experimental demonstrated that proposed models perform consistently better at text mining tasks compared to baseline methods.