Improving Ancient Chinese Word Segmentation With Knowledge‐Enhanced Prompting for Large Language Models

Meng-Tian Tang, Chenggang Mi · International Journal of Intelligent Systems · 2025

This paper introduces a cost‐effective prompt optimization strategy for ancient Chinese word segmentation using large language models, aiming to mitigate the substantial computational resources and training expenses of fine‐tuning. We developed two knowledge‐enhanced frameworks, a General Knowledge Prompt framework and a Domain‐Specific Knowledge Prompt framework, and evaluated their effectiveness across various ancient Chinese corpora using seven mainstream LLMs, including ERNIE Bot, Qwen, SparkDesk, DeepSeek, ChatGPT, Gemini, and Copilot. Our findings confirm that both prompt frameworks enhance the segmentation capability of LLMs to varying extents, with the Domain‐Specific Knowledge Prompt framework yielding the most significant improvements. Notably, the DeepSeek model achieves 94.01% F 1 score (94.24% precision, 93.79% recall) on the test set, while the Qwen model demonstrates a remarkable 15.73% increase in the F 1 score with the Domain‐Specific Knowledge Prompt framework. Our ablation studies indicate that the entries Rules and Examples are the most crucial to the success of prompt frameworks, effectively addressing the challenges of rule inconsistency and insufficient annotated data.

Read the paper · More papers on PaperTik