Efficient Nearest Neighbor Search in Large-Scale Text Data Using Optimized Locality Sensitive Hashing in Distributed Computing Environments

Shreya G. Tendulkar · 2025

Efficient nearest neighbor search is crucial for applications such as web search, mobile browsing, and Natural Language Processing (NLP), where rapid text retrieval from vast datasets is required. As data scales, traditional brute- force approaches become computationally expensive, necessitating optimized methods. This work introduces an enhanced Locality Sensitive Hashing (LSH) framework tailored for distributed computing environments, leveraging Hadoop for large-scale processing. By incorporating optimized LSH variants, our approach improves efficiency while maintaining high recall rates. Experimental evaluations demonstrate that our method surpasses conventional LSH implementations in retrieval accuracy while ensuring computational scalability, making it suitable for real-world applications requiring rapid text similarity assessments.

Read the paper · More papers on PaperTik