An Index Construction and Similarity Retrieval Method Based on Sentence-Bert
Zhimin Wei, Xiaowei Xu, Chenglin Wang, Zhenyu Liu, Peng Xin, Wei Zhang · 2022 7th International Conference on Image, Vision and Computing (ICIVC) · 2022
Semantic Similarity Retrieval is a fundamental task in natural language processing. Instead of sticking to the literal meaning of query, find the vector representation of text by the most advanced semantic index model, index them in high-dimensional vector space, and measure the similarity between query vectors and indexed documents. Basic retrieval approaches can be run on a stand-alone machine, using in-memory algorithms for approximate nearest neighbor search. However, these functions include neither the data management nor the high availability of distributed solutions. This paper designs a method to regenerate semantic dense vector by pre-training a representation model such as Sentence-Bert into string format text and combining it with an inverted index from a Full-text search engine to realize the semantic retrieval function. Solving the problem of text and vector recall must be done in two parts, speeding up retrieval efficiency and taking advantage of robustness and ubiquity. Experimental results show that the effect is better than BM25 and the approximate nearest neighbor search on the LCQMC, OPPO and Quora text similarity datasets.