Disentangled Attention Based Multi-Task Learning with Structural Estimations for Audio Feature Processing and Identification

Aanal Sanjivbhai Raval, Arpita Pareshkumar Maheriya, Shailesh Panchal, Komal Borisagar · 2024

Self-supervised learning utilizes unlabelled data to extract latent representations, replacing conventional input features in downstream tasks. In the realm of audio and speech signal processing, an array of features has been engineered over time through extensive research. Interestingly, predicting these features has emerged as a pertinent pretext task, yielding valuable self-supervised representations that prove effective for subsequent tasks. However, the effective fusion of these pretext tasks to enhance performance in downstream tasks lacks comprehensive exploration and understanding. In this paper, We applied attention base knowledge transfer (ATK) base proposed model with DeBERTa using the specified dataset and proceeded to compare its accuracy against established state-of-the-art models. Subsequently, we evaluated the training, validation, and testing loss of our multitask model in comparison with high-performing existing models. The results demonstrated that our proposed model exhibited the lowest test loss, reaching a remarkable value of 0.018. Moreover, we can extend this research to cover more complex tasks beyond audio data classification with large unlabelled datasets.

Read the paper · More papers on PaperTik