An Exploration Based on the Pre-Trained Model of Literary Knowledge Graph
Xuehua Li · 2025
In today's era, if the field of language and literature is to be deeply integrated with artificial intelligence, it is imperative to build characteristic domain models and evaluation benchmarks. In this paper, Baichuan2-7B-Chat and Qwen-7B-Chat are selected as the base models, and a pretraining model for literary texts based on annotation enhancement is proposed, MythBERT, which improves the hidden language model strategy of BERT, uses a large number of high-quality domain data to train and fine-tune the domain, continuously strengthens the information processing ability of the model's domain knowledge, and can construct a pre-training model in the field of literature at a low cost. This article makes it a truly dedicated AI assistant in the field of language and literature. The purpose of this paper is to promote the cross-integration of NLP technology and the field of language and literature, and to provide a reference for model selection for field applications.