Exploring Human Non-Speech Sound Recognition: Insights from the Nonspeech7K Dataset
CH. V. N. Vaibhav Simha, Ramesh Kumar Bhukya · 2025
Analysis of non-speech sounds produced by humans is an area of speech recognition that has largely not been paid much heed. The lack of proper exhaustive datasets is one of the reasons for the slow-paced research on it, and the inherent difficulty in annotating and collecting clean acoustic data samples has contributed massively to this scarcity of datasets. Recently, a new well-annotated dataset of around 7 thousand human non-speech acoustic samples with 7 classes was introduced—the Nonspeech7k dataset. The dataset is analyzed, and the performance of various machine learning algorithms trained on this dataset has been evaluated. After extracting features such as Mel Frequency Cepstral Coefficients (MFCCs), Mel-spectrograms, Chromagrams, Spectral Contrast, and Tonnetz, the system performance achieved an impressive accuracy of 96.21% using CatBoost, followed by the ExtraTrees classifier with 95.74%. The relevance of this research is in medical applications, particularly in those involving categorization and analysis of non-speech indicators such as coughs, yawns, sneezing, et cetera. This can be further extended to more specialized applications targeting infants through infant cry analysis and also to empower people who are hard of hearing and who cannot express themselves.