Document Storage and Retrieval with Efficient Clustering Using the Vector Space Model

Jyoti Manoorkar, Siddhi Rawool, Shruti Shinde, Nazaat Shaikh, Devansh Thaliya · 2025

The exponential growth of unstructured digital data transcends traditional document-retrieval systems wherein advanced strategies are built that will help realize improved efficiency and scalability. Here is a retrieval system that combines clustering techniques with the Vector Space Model and cosine similarity. The work retrieved in this study integrates k-means and hierarchical clustering, thus reducing search space, allowing enhanced retrieval speed and scalability. The clustering-augmented VSM has better recall and F1-scores especially for semantically complex queries compared to the BM25 model. The use of LSM trees improves indexing and can be used to manage high-dimensional data efficiently. Empirical tests confirm the effectiveness of the system, making it ideal for healthcare, finance, and legal research, where real-time and precise data retrieval is important.

Read the paper · More papers on PaperTik