Self-supervised Rewiring of Pre-trained Speech Encoders:Towards Faster Fine-tuning with Less Labels in Speech Processing

Hao Yang, Jinming Zhao, Gholamreza Haffari, Ehsan Shareghi · 2022

Pre-trained speech encoders have facilitated great success across various speech processing tasks.However, fine-tuning these encoders for downstream tasks require sufficiently large training data to converge or to achieve stateof-the-art.In text domain this has been partly attributed to sub-optimality of the representation space in pre-trained Transformers.In this work, we take a sober look into pre-trained speech encoders and rewire their representation space without requiring any task-specific labels.Our method utilises neutrally synthesised version of audio inputs along with frame masking to construct positive pairs for contrastive self-supervised learning.When it is used for augmenting the WAV2VEC 2 encoder, we observe consistent improvement of isotropy in the representation space.Our experiments on 6 speech processing tasks, exhibit a significant convergence speedup during task fine-tuning as well as consistent task improvement, specially in low-resource settings.1

Read the paper · More papers on PaperTik