A Method for Modeling the Neural Network for Term Extraction Based on Bilingual Sentence Alignment Corpus
Baosheng Yin, Binbin Zhang, Shaoming Li, Meishu Zhao · 2022 IEEE 2nd International Conference on Power, Electronics and Computer Applications (ICPECA) · 2022
The implementation of the current term extraction model is mainly dependent on the training on manual tagging of a bilingual term corpus. It will take a great amount of manpower to tag terms in a massive corpus. In order to solve this problem, this paper proposes a method for modeling the neural network for term extraction based on bilingual sentence alignment corpus, which implies that GIZA++ is used to tag terms in terms of bilingual word alignment based on bilingual sentence alignment corpus, thus generating the tagged corpus for bilingual word alignment as the initial training input of neural network. In order to fully explore the potential language features between sentences, from the perspective of deep learning, an unsupervised training algorithm for word alignment based on Word2vec training word vector and integrated into the neural network was proposed. The experiment was carried out based on English-Chinese patent and standard sentence alignment corpus, with an accuracy of 0.8662.