On Using Pre-Trained Embeddings for Detecting Anomalous Sounds with Limited Training Data
Kevin Wilkinghoff, Fabian Fritz · 2023
Using embeddings pre-trained on large datasets as input representations is a popular approach for classifying audio data in case only a few training samples are available. However, for anomalous sound detection pre-trained embeddings usually perform worse than directly training a model because subtle changes indicating anomalous data are not captured sufficiently well. In this paper, the potential of using pre-trained embeddings for detecting anomalous sounds with limited training data is investigated. In experiments conducted on datasets for anomalous sound detection with domain shifts and few-shot openset classification, it is shown that with increasing openness directly training a model on the original data leads to better performance than using pre-trained embbedings as input. Regardless of the input representation, the presented system achieves a new state-of-the-art performance for few-shot open-set classification in all pre-defined openness settings and is made publicly available.