MIGUE-Sim: Speeding Up Similarity Queries with Native RDBMS Resources
Igor Alberte Rodrigues Eleutério, Mirela T. Cazzolato, Larissa Roberta Teixeira, Marco Antonio Gutierrez, Agma J. M. Traina, Caetano Traina · 2024
Many applications require storing, managing, and retrieving complex data, such as multidimensional vectors and images in databases. In this paper, we propose MIGUE-Sim, a system to quickly execute exact Range and kNN similarity queries in Postgres. The queries are expressed following a straightforward, SQL-compatible representation seamlessly integrated into the language, whereas the system executes each query using just the native resources of Postgres. MIGUE-Sim uses the Postgres's Cube native extension to perform kNN faster, using the GIST R-Tree index available. The execution of kNN in our system without any index overcame our main competitor by up to 10% in execution time. However, when using GIST R-Tree, MIGUE-Sim can significantly speed up queries - experiments revealed that MIGUE-Sim is up to 96% faster than our closest competitor. We contribute with a framework that uses existing index structures from Postgres to speed up kNN queries with no modifications on the RDBMS; with the evaluation of different ways to write similarity queries in plain SQL; with the implementation of kNN queries with indices already available on Postgres. Our approach is easy to understand, to use, and it is extensible to include other distance functions.