Weakly Supervised Text-to-SQL Parsing through Question Decomposition

Tomer Wolfson, Daniel Deutch, Jonathan Berant · Findings of the Association for Computational Linguistics: NAACL 2022 · 2022

Text-to-SQL parsers are crucial in enabling non-experts to effortlessly query relational data.Training such parsers, by contrast, generally requires expertise in annotating natural language (NL) utterances with corresponding SQL queries.In this work, we propose a weak supervision approach for training text-to-SQL parsers.We take advantage of the recently proposed question meaning representation called QDMR, an intermediate between NL and formal query languages.Given questions, their QDMR structures (annotated by non-experts or automatically predicted), and the answers, we are able to automatically synthesize SQL queries that are used to train text-to-SQL models.We test our approach by experimenting on five benchmark datasets.Our results show that the weakly supervised models perform competitively with those trained on annotated NL-SQL data.Overall, we effectively train text-to-SQL parsers, while using zero SQL annotations.Question Decomposition: 1. papers 2. #1 in PVLDB 3. authors of #2 4. number of #2 for each #3 5. #3 where #4 is more than 10 1 SQL Synthesis Question: "Which authors have more than 10 papers in the PVLDB journal?" Weak Supervision Execution-guided SQL candidate search 2 Training a Text-to-SQL Model 3 Question: "Which authors have more than 10 papers in the PVLDB journal?

Read the paper · More papers on PaperTik