Feature Extraction Strategy of Translated Text Based on Neural Network

Han Gong, Xiaopei Wang · 2023

In order to better solve the problems of difficult processing of information in text semantics, too much artificial prior information, and unstable performance in traditional natural language processing, a text feature extraction method based on neural network was proposed on the basis of relevant theories of neural network. This method constructs a translation model based on neural networks and an improved deep neural network feature extraction strategy. In order to enable the deep learning model to better handle the semantic content of translation, this paper first designs a semantic based translation segmentation model, which improves that the existing translation segmentation algorithm is not suitable for single semantic segmentation, and this model can also model language in training, and use the deep learning neural network to provide high-quality translation semantic word vectors for other natural language processing systems. In addition, we conducted a separate word segmentation test for sentences with semantic ambiguity, and then analyzed the existing hidden Markov model, maximum entropy Markov model, conditional random field model and other models based on traditional machine learning algorithms, as well as the advantages and disadvantages of deep learning algorithms such as LSTM sequence tagging model, LSTM-CRF sequence tagging model and Bi LSTM-CRF model. Through algorithm training, the proposed improved algorithm model iterated over 21 corpora. The result finding that the single stack bi LSTM CRF model constructed by a single-layer LSTM neural network was chosen as the final model in this paper. A total of 21 iterations of the corpus were trained, and the accuracy of all labels was 98.70%. On the development set, the recall rate is 89.31%, with a peak F score of 93.88. In addition, the accuracy of the test set obtained in the experiment was 97.81%, the recall rate was 83.94%, and the F-value was 90.35 points. This article improves the neural network structure in Bi-LSTM-CRF and obtains the performance of existing models on the named entity recognition test set.

Read the paper · More papers on PaperTik