Reusing Speech Techniques for Video Semantic Indexing [Applications Corner]
Koichi Shinoda, Nakamasa Inoue · IEEE Signal Processing Magazine · 2013
Many techniques developed in speech research have been successfully employed in other fields, such as automatic video semantic indexing. In this application, a user submits a textual input query for an desired object or a scene to a search system, which returns video shots that include the object or scene. Recently, a new method using Gaussian-mixture model (GMM) supervectors and support vector machines (SVMs) was proven to be very effective. In this method, speech technology such as speaker verification and adaptation techniques play very important roles.