Content Sifting Storage: Achieving Fast Read for Large-scale Image Dataset Analysis
Yu Liu, Hong Jiang, Yangtao Wang, Ke Zhou, Yifei Liu, Li Liu · 2020
Analyzing large-scale image dataset requires all images to be read from disks first, leading to high read latency. Therefore, we propose a Content Sifting Storage (CSS) system, which aims to reduce the read latency by only reading sifted relevant data. CSS generates embedded content metadata via deep learning and manages the metadata via Semantic Hamming Graph, which achieves fast read based on content similarity meeting the given analysis. Extensive experimental results on image datasets show that compared with conventional semantic storage systems, our CSS can greatly reduce the read latency by 82.21% to 94.8% with more than 98% recall rate.