Research on Multi-Turn Chinese NL2SQL Methods Based on Semantic Rewriting
<p>Shuaike Guo, Yongzheng Yang</p> · Academic Journal of Computing & Information Science · 2025
In the era of data intelligence, Natural Language to SQL (NL2SQL) technology serves as a core interface for human-machine data interaction, making its performance optimization in multi-turn Chinese dialogue scenarios highly valuable for research. In addressing the issue where reference resolution and semantic omission in the Chinese context can lead to gaps in understanding user intent, this paper proposes a model, RW-T5, integrated with a semantic rewriting mechanism. This model is based on the pre-trained T5 architecture and utilizes hierarchical modeling of dialogue history along with turn-aware encoding to accurately parse semantic unit segmentation and temporal dependencies in multi-turn interactions. It features an innovative design for a global context injection and bidirectional cross-attention fusion module, enabling the capture of both the overall semantic focus and fine-grained word-level semantic details. Utilizing a sequence optimization strategy based on multi-dimensional semantic feature fusion, the model effectively performs explicit resolution of implicit reference relationships and logical completion of omitted semantics in multi-turn dialogues, providing a semantically complete and structurally standardized input for subsequent SQL statement generation. Experimental validation on the large-scale Chinese multi-turn dialogue benchmark dataset, CHASE, shows that this model significantly outperforms other advanced NL2SQL parsing methods, fully validating the effectiveness of the dynamic semantic rewriting mechanism and hierarchical modeling approach, and offering an effective solution for the engineering implementation of intelligent data interaction systems in Chinese multi-turn dialogue scenarios.