Text Similarity Based on Siamese Multi-Output Bidirectional Long Short-Term Memory
Zhengfang He, Cristina E. Dumdumaya · 2024
Text similarity measurement has become even more important in natural language processing. Usually, the output of the BiLSTM only takes the last vectors of forward and reverse sequences. However, when the input sentence sequence is very long, the BiLSTM easily forgets the words at the beginning. Hence, the output of the BiLSTM in this paper not only takes the last vectors of forward and reverse sequences but also includes the half vectors. With the addition of the half vectors, the model is termed Multi-Output BiLSTM (MOBiLSTM). Since this paper focuses on text similarity, it uses the Siamese architecture; therefore, the proposed model is named S-MOBiLSTM. Furthermore, the experiments use the STS-B dataset, with Bidirectional Encoder Representations from Transformers (BERT) as the pretraining model. The S-MOBiLSTM model achieved 88.5% Spearman's rank correlation coefficient, which increased by 2.0%, compared with the fine-tuned BERT Large model. The experimental results show that the S-MOBiLSTM model performs better than other algorithms. Finally, this paper presents a general application framework based on the text similarity algorithm.