An Optimization Method TSDAE -based for Unsupervised Sentence Embedding Learning
YueYing Zhu, Xu Zhang, Qi Zhao · 2024
As for the research on learning sentence embeddedness, most of the research methods involved at present are designed for labeled data in the general field. The evaluation is also basically carried out on a single task Semantic Textual Similarity (STS). However, many tasks in specific fields are involved in the real society, which presents certain challenges compared with general fields. For example, there are most unlabeled data sets and certain professional knowledge. The learning methods in general fields are difficult to adapt to specific fields or tasks. To solve the problem of unlabeled data sets, specialized knowledge in specific fields or tasks, we propose an optimization method TSDAE-based for unsupervised sentence embedding learning(AOM), to optimize and improve the existing unsupervised sentence embedding learning based on pre-trained transformers and sequential denoising auto encoder (TSDAE)[9]. We firmly believe that improving the training difficulty of noise reduction autoencoders in unsupervised sentence representation learning is the key to achieve further breakthroughs in semantic text similarity (STS) tasks. Based on this idea, this study innovates the existing technology. Our research results show that in STS (Semantic Text similarity) downstream task testing, our proposed improvement strategy only needs one-tenth of the corpus data of the original method to exceed the performance of the original model.