Chimera-VDB: Mixed-Precision Vector Database with HNSW Index for RAG-LLM
Naoshi Yamane, Michael Ryan Zielewski, Takaki Nakamura, Takuo Suganuma · 2025
In recent years, vector databases have become a core component in Retrieval-Augmented Generation (RAG) systems for Large Language Models (LLM), enabling fast retrieval of documents similar to a given query. However, storing a large number of high-dimensional vectors requires substantial storage capacity. A common solution is to reduce vector precision through quantization, but this often degrades retrieval recall. To address this trade-off, we propose Chimera-VDB, which uses a mixture of high- and low-precision vectors to reduce data size while maintaining retrieval recall. Our approach leverages the graph structure of Hierarchical Navigable Small World (HNSW) networks, selectively preserving only the most search-critical vectors in high-precision, while quantizing the rest to lower precision. Experimental results show that our method reduces storage usage to 24% compared to storing all vectors in FP32, while maintaining 93% recall, demonstrating its effectiveness in balancing storage capacity and retrieval performance.