Research on Key Methods for Extracting High-Quality Chinese Corpus Based on Common Crawl

Leling Xiao, Zhigang Zhao, Chunxiao Wang, Jian Zhang, Fulai Liu · 2025

With the rapid development of large-scale Chinese pre-trained language models and natural language processing technologies, the demand for large, high-quality Chinese corpora is significantly increasing. However, current Chinese corpora often fail to meet the demands of high-quality training due to their limited size and substandard quality. This paper introduces a method to extract high-quality Chinese corpora based on the Common Crawl dataset, enhancing existing quality filtering and deduplication methods. In terms of deduplication, this paper enhances the Simhash algorithm to tackle the issue of feature word co-occurrence effectively and introduces a deduplication method based on sentence clusters. Additionally, this paper proposes an offensive speech detection method, ODet, based on prefix tuning and prompt learning. ODet has demonstrated superior performance compared to existing offensive speech detection methods and has been successfully applied in the corpus processing workflow. We successfully extracted a 500GB high-quality Chinese corpus, CHCorpus, employing this approach. Furthermore, we applied a perplexity calculation method to segment CHCorpus into head, middle, and tail sections. The experimental outcomes indicate that CHCorpus excels in perplexity assessments, several CLUE benchmark tasks, and text summarization, surpassing Wikipedia, CLUECorpus2020, and CCNet corpora. This research provides a robust data foundation for training large-scale Chinese pre-trained models, significantly improving their robustness and security.

Read the paper · More papers on PaperTik