Web-scale system for image similarity search: When the dreams are coming true

David Novák, Michal Batko, Pavel Zezula · 2008

Digital images have become a commodity which is searched on the Web as ordinarily as web pages. However, current large-scale engines search the images only on the basis of their annotations, while the content-based similarity systems do not seem to be ready for such scales. In this paper, we open the way to Web-scale image similarity search. We present a flexible system based on the metric space model and on the peer-to-peer paradigm. It uses M-Chord and M-Tree structures as its fundamental components and measures the image similarity by a combination of five MPEG-7 features. The system has been implemented including a graphical interface for online demonstrations and it currently indexes 10 million images crawled from the Web. We propose a novel strategy for approximate evaluation of similarity queries and we test its performance by a series of experiments. The results show that the system provides high-quality answers with response times around 0.5 second.

Read the paper · More papers on PaperTik