On-Edge Device Optimization Using Multiple Classification Method for a Cat and Dog Audio Classifier

Wicaksono Leksono Muhamad, Dwi Ahmad Dzulhijjah, Fathin Difa Robbani, Dyah Aruming Tyas · 2025

The Internet of Things (IoT) relies on edge devices with limited memory, power, and computing resources, making machine learning (ML) implementation challenging. Recent advances in model quantization have enabled the deployment of reduced-precision ML models. This paper explores deploying both neural and traditional ML models on the Arduino Nano 33 BLE Sense, a Cortex-M processor with 256 KB RAM and 1 MB ROM. We compare models for classifying cat and dog audio recordings from a Kaggle dataset. Our approaches include Mel-Frequency Cepstral Coefficients (MFCC), Mel-Frequency Energy (MFE), and a spectrogrambased preprocessor combined with a Convolutional Neural Network (CNN). We also extract features like Chroma, Spectral Contrast, Spectral Centroid, Zero Crossing Rate, and Tempo, computing their statistical summaries to train classifiers such as Support Vector Machines (SVM), Random Forests (RF), and Gradient Boosting Machine (GBM). Among the models, the MFE-based CNN achieved the highest accuracy of 95 2%, while the MFCC-based CNN had the fastest inference time of 3.00 ms and the lowest RAM usage of 3.80 KB. The SVM model was most efficient in flash memory usage, requiring only 9.72 KB. These results highlight the trade-offs between accuracy, speed, and resource consumption when deploying ML models on resource-constrained IoT devices. This study demonstrates the feasibility of deploying ML models on such devices for effective audio classification and provides information on optimal model selection based on performance metrics and hardware limitations.

Read the paper · More papers on PaperTik