MIGUE-Sim: Speeding Up Similarity Queries with Native RDBMS Resources

Igor Alberte Rodrigues Eleutério, Mirela T. Cazzolato, Larissa Roberta Teixeira, Marco Antonio Gutierrez, Agma J. M. Traina, Caetano Traina · 2024

Many applications require storing, managing, and retrieving complex data, such as multidimensional vectors and images in databases. In this paper, we propose MIGUE-Sim, a system to quickly execute exact Range and kNN similarity queries in Postgres. The queries are expressed following a straightforward, SQL-compatible representation seamlessly integrated into the language, whereas the system executes each query using just the native resources of Postgres. MIGUE-Sim uses the Postgres's Cube native extension to perform kNN faster, using the GIST R-Tree index available. The execution of kNN in our system without any index overcame our main competitor by up to 10% in execution time. However, when using GIST R-Tree, MIGUE-Sim can significantly speed up queries - experiments revealed that MIGUE-Sim is up to 96% faster than our closest competitor. We contribute with a framework that uses existing index structures from Postgres to speed up kNN queries with no modifications on the RDBMS; with the evaluation of different ways to write similarity queries in plain SQL; with the implementation of kNN queries with indices already available on Postgres. Our approach is easy to understand, to use, and it is extensible to include other distance functions.

Read the paper · More papers on PaperTik