Audio Synthesis-based Data Augmentation Considering Audio Event Class

Toki Sugiura, Akio Kobayashi, Takehito Utsuro, Hiromitsu Nishizaki · 2021 IEEE 10th Global Conference on Consumer Electronics (GCCE) · 2021

Although many studies have been conducted on environmental audio detection or classification, most have been done using well-developed corpora. Build a training corpus is necessary to perform new audio event detection tasks using deep learning. However, the cost of data collection and labeling is very expensive. Research has focused on improving model accuracy by efficiently increasing the amount of training data using data augmentation approaches on the collected data. Data augmentation methods for signals, such as audio, have not yet been established, so this study proposes a data augmentation method based on audio synthesis using audio similar to those used in the audio detection or classification task. The proposed method performs data augmentation using appropriate seed audio corresponding to the target sound class to be recognized. The experimental results showed that the proposed method improved the F1 score by 12.9 points compared with the general augmentation method in the audio tagging task in baseball games.

Read the paper · More papers on PaperTik