Word and Text Similarity Using Classical Word Embeddings in Quantum NLP Systems

Damir Ćavar, Koushik Reddy Parukola · 2025

Word embeddings and vector representations for linguistic units are important in Natural Language Processing (NLP) and AI systems. Engineering of such models involves significant effort, large amounts of data, and costly computations. Many models and systems have been released to the general public, enabling significant research and development in the classical NLP domain. We demonstrate how word and text embeddings as n-dimensional dense vectors of n real numbers in classical computers can be mapped to highly compressed quantum states or quantum embedding using log(n) qubits. In similarity experiments performed on classical and quantum computers, we show no significant information loss that would affect vector distance scores representing words’ semantic similarity. The results show that existing NLP embedding models in classical computing environments can be used in quantum computing. We discuss the issues and limitations of this approach in the context of current quantum computing environments.

Read the paper · More papers on PaperTik