Anchor Your Embeddings Through the Storm: Mitigating Instance-to-Document Semantic Gap

Renato Marques Barros, Kristen K. Arguello, Jônatas Wehrmann · 2024

Large Language Models (LLMs) have revolutionized the field of natural language processing with their remarkable ability to generate coherent responses. Despite their impressive capabilities, LLMs grapple with the challenges of learning from private data while ensuring the relevance and timeliness of the information they provide. To overcome this challenge, retrieval-based strategies for generation have become essential, though they require clean data and a complex indexing pipeline. In this paper, we introduce Anchor Embeddings, a novel embedding enhancing technique designed to mitigate the instance-to-document semantic gap that often hinders the retrieval process. Our method proposes the usage of an anchor embedding to serve as a semantic beacon, adding more context to smaller text segments extracted from the same source. This extra embedding acts as a holistic representation regarding the original document that is merged into the instance embeddings. It can be derived based on a diversity of strategies that involve semantic representation and localization cues. Our empirical analysis on the MS MARCO dataset reveals that Anchor Embeddings can significantly outperform traditional retrieval methods, boasting up to an 14% performance improvement. The elegance of our approach lies in its simplicity and robustness, providing more specific context while maintaining the same time complexity for retrieval of the baseline approach.

Read the paper · More papers on PaperTik