Hybrid Self-Interactive Attentive Siamese Network for Medical Textual Semantic Similarity

Hongda An, Di Jia Wu, Zhengguang Li · 2020

The rapid development of medicine produces a large number of medical texts, but it is difficult to process these texts due to many similar sentences. Therefore, estimating the similarity of medical texts has become a key technology, filtering out medical texts quickly. Nowadays, many methods for estimation similarity between medical sentences extract semantic features mainly via Siamese network. However; these methods don't achieve the best results due to the large amount of noise in the texts. To improve the performance of the Siamese network, a hybrid self-interactive attention model is proposed in this paper. The aim is to reduce the noise of the text and strengthen the token with high correlation between the two texts. In addition, this proposed model also uses BERT as the embedding layer to carry out a preliminary pre-training of text. Then, two datasets are employed to verify the effectiveness of our method and our method achieves better results on Pearson correlation coefficient, compared with the other existing methods. The experimental results still indicate that the results of pre-trained BERT are better than that of Word2Vec, and the hybrid self-interactive attention model obtains better results due to the effect of interactive attention.

Read the paper · More papers on PaperTik