Enhancing Transformer-Based Semantic Search by Integrating Human-Crafted Knowledge Into Sentence Embeddings
Zachary Weinfeld, Lubomir Stanchev · 2025
Semantic search is the problem of document retrieval based on the meaning of the words and phrases in the query and documents and not solely based on word commonality. A recent approach to semantic search involves using sentence transformers to generate embeddings for both the search queries and documents, allowing for the comparison of their vectors. While many different embedding models are widely used, our approach integrates these models with human-crafted knowledge in a novel way, resulting in an improvement in the Mean Average Precision (MAP) scores. Embedding models are usually trained without user input by using a corpus of documents and remembering what words and phrases follow each other. Our approach enhances these techniques by performing query and document expansion using similar words based on information from WordNet. In our experiments, we use 23 different sentence embedding models on the Cranfield benchmark [1] and achieve a statistically significant improvement in MAP scores, with a p-value of 0.047. The average MAP improvement over all models is a little above 2%.