Sequential Feature Augmentation for Robust Text-to-SQL

Hao Shen, Ran Shen, Gang Sun, Yiling Li, Yifan Wang, Pengcheng Zhang · 2023

The task of converting natural language queries into SQL queries, known as Text-to-SQL, plays a crucial role in bridging the gap between human language and database systems. However, Text-to-SQL systems face numerous chal-lenges due to the inherent complexity of natural language understanding and the structural differences between text and SQL. On the one hand, the accuracy of SQL generation of previous Text-to-SQL models needs to be further improved. On the other hand, the robustness of previous Text-to-SQL models is not enough, so the usability is not strong enough. In this paper, we propose a novel approach called Sequential Feature Augmentation Method (SFAM) to enhance the robustness of Text-to-SQL models. In SFAM, we generate a corresponding augmented feature sequence for each preliminary feature sequence in the text encoding stage. To strengthen the ability of Text-to-SQL model to distinguish the original sequence features and the augmented sequence features, we designed a method based on sequential consistency learning. In addition, we design a pooling method based on attention mechanism and mutual information maximization to reduce the loss of information during pooling. Constrained by these methods during training, our model not only greatly increases robustness but also achieves better accuracy of SQL generation. We construct a dataset based on the marketing business data of the State Grid Corporation of China. Experiments on this dataset demonstrate our SFAM can achieve state-of-the-art effects on the Text-to-SQL task. Besides, to verify the robustness of our model, we randomly sample some data in the test set of the above dataset, and generate synthetic queries based on ChatGPT. Experiments on this test set verify the robustness of our SFAM.

Read the paper · More papers on PaperTik