Combining Document Embedding Techniques for Clustering and Analysis of Extractive Summaries
Leena Sinha, S. Jaya Nirmala · 2021 IEEE 8th Uttar Pradesh Section International Conference on Electrical, Electronics and Computer Engineering (UPCON) · 2021
In today’s fast-paced world, where we are flooded with massive amounts of unstructured and written data every second, a streamlined approach to make these documents readable is required. We need to cluster these documents in such a way that documents with comparable contexts and categories are automatically grouped. Therefore, it is of interest to understand and study how various algorithms in existence merge together to provide us with a suitable approach to group subjects of similar nature. This work aims at stacking different algorithms to propose an alignment to find out which stack works best for the given use case. Here documents of different categories have been utilized, which undergoes different embedding techniques along with dimensionality reduction, which are later clustered using k-means clustering, to compare which stack works best. The dataset contains news article from five different categories namely entertainment, sports, business, politics, and technology. We can observe how well the different embedding algorithms preserve semanticity and context of the given document. The performance of different stacks in each category of the documents has been analyzed using clustering and the best results have been observed from TF-IDF and word2vec embeddings.