A text restoration model for ancient texts based on pre-trained language model RoBERTa

Zhongyu Gu, Yanzhi Guan, Shuai Zhang · 2024

We propose an Ancient Chinese Text Restoration Model (ACTRM) that leverages the pre-trained RoBERTa model. This model is designed to continuously assimilate semantic information from the surrounding context to reconstruct missing text segments. The model's predictive capabilities are notably enhanced through the incorporation of phrase-level semantics within the input embedding layer and the application of a KAN network in the prediction layer. The primary objective of this model is to augment the workflow of specialists engaged in the restoration of ancient texts. To facilitate the training of the model, a dataset comprising ancient Chinese texts across diverse subjects—including science, technology, geography, history, and science fiction—has been developed. The model achieves an accuracy rate of 90.6% for Top-5 predictions and 76.7% for individual word predictions within this dataset, while also ensuring semantic coherence and fluency in the reconstructed text. These findings substantiate the model's efficacy as a supportive tool for the restoration of ancient texts, potentially enhancing both the efficiency and accuracy of such restoration efforts.

Read the paper · More papers on PaperTik