Research on the Construction Method of Chinese - Vietnamese Parallel Corpus
Shiying Tu, Haojin Hu, Ronglyu Sun, Yanmei Jing, Wenxue He · 2019
The Chinese-Vietnameseparallel corpus is the basic research problem in the fields of natural language processing. The traditional methods use the DOM tree or element anchors in HTML extract parallel sentences with low accuracy and slow alignment speed. Therefore, this paper proposes a new Web-based Chinese-Vietnamese parallel corpus construction scheme. The scheme will determine the parallel web page through the LDA (Latent Dirichlet Allocation) and Gibbs Sampling. And the BeautifulSoup and regular expression will be used to crawl the webpage text and clean the corpus. The DOM tree and the element anchors in HTML are used to optimize the extraction of parallel sentence pairs. Combined with the sentence length and Champollion algorithm, the dynamic programming algorithm is adopted to improve the correct rate and recall rate of sentence alignment. The program successfully established a million-level Chinese-Vietnamese parallel corpus.