An acoustic sensing system for noise monitoring and source identification using transfer learning
Dolvara Gunatilaka, Wudhichart Sawangphol, Thanakorn Charoenritthitham, Thanawat Kanjanapoo, Teerapat Burasotikul, Kittikawin Pongprasit · Expert Systems with Applications · 2025
• A scalable and cost-effective IoT-based noise monitoring system is developed for smart urban environments. • The system integrates low-cost acoustic sensors, edge-based CNN noise classification, and a web interface for data visualization. • The proposed model, using Mel Spectrogram and MobileNet, achieves a 90.18 % classification accuracy. • The deployed system demonstrates low latency, with 0.37 s transmission delay and 2.5 s processing time on a Raspberry Pi. Increasing noise pollution in urban areas underscores the need for an autonomous system to monitor and control noise. Beyond detecting noise levels, identifying noise sources further improves noise management. This work presents a scalable IoT-based sensing platform for smart environment applications. The system integrates low-cost devices for acoustic measurement, edge devices to enable noise source identification, a back-end infrastructure crucial for efficient acoustic data and device management, and a web-based application facilitating noise data visualization. Our study explores three feature extraction techniques and eight Convolutional Neural Network (CNN)-based pre-trained models for noise classification on the resource-constrained Raspberry Pi platform and compares their performance. Leveraging pre-trained models helps speed up the model development process. UrbanSound8k, ESC-50 datasets, and audio data collected with our low-cost microphone are used for model development and validation. The evaluation results show that our hierarchical model, utilizing the Mel Spectrogram feature extraction method and a MobileNet model, achieves the highest accuracy of 90.18 %. Furthermore, we deploy the system and assess its performance. Our system can reliably transmit audio data with an average delay of 0.37 s, and the Raspberry Pi can perform feature extraction and classification within an average of 2.5 s. Hence, our solution offers a comprehensive and cost-effective solution to enhance noise management and control.