Speech classification by using binary quantized SIFT features of signal spectrogram images

Bui The Duy, Nguyễn Quang Trung · 2016

SIFT features are not only widely used in content based image retrieval, but also used in speech perception problem by using SIFT features combination with LNBNN classifier. With the combination of SIFT and LNBNN, speech perception problem is quite highly accurate. In this approach, each speech signal is converted into an spectrogram image, then SIFT features are extracted from this image. Typically, a few thousand key-points are extracted from each image spectrogram. This approach retains all discriminative features of speech signals. However, the big limitation of this approach is that the need of memory to store all SIFT features. In order to reduce the need of memory, we quantize SIFT features from 128 bytes to 128 bits and encode them into 16 bytes for each SIFT features. We show that SIFT features perform surprisingly well even after quantizing each component to binary, when the medians are used as the quantization thresholds. Quantized features preserve both distinctiveness and matching properties. The LNBNN perform searching K-nearest neighbor in classification, thus it uses the KD-TREE structure to speed up searching in classification, however, KD-TREE cannot perform with binary SIFT. Therefore, we propose using a multi index hashing to improve the KNN search problem.

Read the paper · More papers on PaperTik