PAUQ: Text-to-SQL in Russian
Daria Bakshandaeva, Oleg Somov, Ekaterina V. Dmitrieva, Vera Davydova, Elena Viktorovna Tutubalina · 2022
Semantic parsing is an important task that allows to democratize human-computer interaction.One of the most popular text-to-SQL datasets with complex and diverse natural language (NL) questions and SQL queries is Spider.We construct and complement a Spider dataset for the Russian language, thus creating the first publicly available text-to-SQL dataset in Russian.While examining dataset components-NL questions, SQL queries, and database content-we identify limitations of the existing database structure, fill out missing values for tables and add new requests for underrepresented categories.We select thirty functional test sets with different features for evaluating the capabilities of neural models.To conduct the experiment, we adapt baseline models RAT-SQL and BRIDGE and provide in-depth query component analysis.Both models demonstrate strong single language results and improved accuracy with multilingual training on the target language.In this work, we also study tradeoffs between automatically translated and manually created NL queries.At present, Russian text-to-SQL is lacking in datasets as well as trained models, and we view this work as an important step toward filling this gap.