Composite Deep Learning Model with Augmented Features for Accurate Animal Sound Detection and Classification

Dilip Singh Sisodia, Mihir Kumar Singh, Ishaan Singhal · 2024

Accurate animal sound detection and classification may assist in decision-making for many applications such as monitoring the range shift of animals due to climate change, biodiversity assessment of an area, and alerting the nearby people to avoid human-animal conflicts. In this study, a composite deep-learning model for classifying and detecting animal sounds is proposed. The model combines bidirectional long short-term memory (LSTM) and sequential convolutional neural network (CNN). The proposed model consists of four convolution layers with max-pooling, two dropout layers, a bidirectional LSTM, and finally, a fully connected layer. The model is evaluated using a dataset consisting of 885 sound clips of 11 different wild and pet animals. To ensure robust training of the composite deep learning model and make it uniform to the maximum possible animal sound disturbances, the used dataset is augmented for feature extraction. The dataset is augmented by adding noise, stretching, rolling, pitch shifting, etc. Several features are extracted from the augmented dataset, including Chroma, Short-Time Fourier Transform (STFT), Mel Spectrogram, Spectral Contrast, Mel frequency cepstral coefficient (MFCC), and Tonnetz. Extracted features are passed as input for training, and model parameters are optimized for better performance. The performance of the proposed model is compared with conventional CNN and bidirectional LSTM separately using various optimizers. CNN and Bidirectional LSTM achieved the highest accuracy of 91.92% and 93.64%, respectively, with Adam optimizer. Simultaneously, the proposed model achieved the highest accuracy of 95.18 % with the Adagrad optimizer.

Read the paper · More papers on PaperTik