Identification of Urban Sounds with Haptic Feedback using Raspberry Pi and LSTM-SVM
Kyle Kenshin T. Morales, Carlo Castillo, Rosemarie V. Pellegrino · 2024
Sound is a physical phenomenon in the form of vibration that contains information. This information is perceived by humans using auditory senses to contextualize their surrounding environment. There are different kinds of sound in which several studies utilized image processing as key feature information, such as spectrogram and MFCC, to classify them using deep learning or machine learning. This study focuses on non-image audio features to classify 5 different audio categories, namely, siren, bike bell, scream, car engine, and street music. The study focuses on the LSTM's ability to contextualize temporal data and SVM's overall ability to generalize information. With LSTM and SVM, the study was able to use Raspberry Pi 4B to run the model and display the prediction on LCD Display, utilizing microphone as the input, and vibration module as a haptic feedback output. The LSTM-SVM model produced from this study was able to output an 84.80% accuracy where the SVM is supported by the activated neurons of LSTM.