Softening quantization in bag-of-audio-words
Stephanie Pancoast, Murat Akbacak · 2014
The audio component of multimedia data can be crucial for multimedia content analysis. Bag-of-audio-words (BoAW) approach is one of the most frequently used methods to represent audio content in multimedia event detection and related tasks. The method, however, has numerous criticisms, amongst which is the loss of information in the “vector quantization” step which generates word-like units. In this work, we address this issue by employing a soft quantization representation where the distance to the nearest codeword is incorporated into the model, rather than only using the nearest codeword's index as is the case with hard quantization. We explore two techniques for soft quantization and apply it to the BoAW for multimedia event detection. We find the best setup yields a 13% improvement in mean average precision, improving performance for 27 of the 30 video events.