Lightweight CNN for Everyday Human Sound Classification on Edge Devices

Venkata Raghava Shashank Viswanathuni, Sarada Puranapanda, Bethi Pardhasaradhi, Sagar Koorapati, Linga Reddy Cenkeramaddi · 2025

This study presents a lightweight Convolutional Neural Network (CNN) that recognizes everyday sounds on IoT edge devices. Audio signals encompassing daily activities are captured via a microphone and transformed into time-frequency images using Continuous Wavelet Transform (CWT). The proposed CNN is optimized for low computational requirements and performs better than established pre-trained models such as DenseNet, EfficientNet, InceptionNet, MobileNet, NASNet, ResNet, VGGNet, and Xception. The model achieves an overall validation accuracy of 84 % in a ten-fold cross-validation framework, with a compact model size of only 854KB and an inference time of 7.3 milliseconds (ms) on a Raspberry Pi. In comparison, EfficientNetV2B1, the second-best performer, achieves comparable accuracy but with a significantly larger size of 27.1 MB and an inference time of 119.99 ms. The model’s inference time was evaluated on various devices, including Raspberry Pi, NVIDIA A100, Tesla V100, and Intel Xeon Platinum, demonstrating its efficiency across different platforms. The dataset includes everyday human sounds, such as crying, applause, breathing, coughing, footsteps, laughter, brushing teeth, snoring, and drinking water. The efficiency and accuracy of the proposed CNN make it ideal for real-time sound recognition in smart home automation, health monitoring, and ambient assisted living, where low latency and resource efficiency are critical.

Read the paper · More papers on PaperTik