Automatic SQL Query Generation from Code Switched Natural Language Questions on Electronic Medical Records
Haodi Zhang, Jinyin Nie, Zeming Liu, Dong Lei, Yuanfeng Song · 2024
Electronic Medical Records (EMRs) meticulously document patient information in relational databases, presenting a challenging task for medical professionals to effectively retrieve this data. Natural Language Question to SQL query (NL2SQL), a critical task in natural language processing (NLP), shows promising performance in addressing this challenge. However, existing medical NL2SQL studies often focus on generating SQL queries from monolingual questions in English. None of them have studied medical NL2SQL for the Chinese domain.In Chinese EMRs, the questions are primarily in Chinese, while many medical terms, such as drugs and diseases, are commonly described in English. This phenomenon is generally known as code-switching (CS) in the field. Given that underlying systems are typically monolingual, CS has proven to pose significant accuracy challenges, as observed in many tasks such as Automatic Speech Recognition (ASR) and Machine Translation (MT). However, the potential effects on EMRs NL2SQL have not been explored. In this study, we investigate the CS-NL2SQL problem, focusing on CS in the context of Chinese questions for the NL2SQL task on EMRs. To assess the model’s performance on this task, we construct the first CS-NL2SQL dataset named CS-MIMICSQL in medical domains. We then explore different CS-NL2SQL architectures along two dimensions: cascaded (translation followed by SQL generation) vs. end-to-end. Our results demonstrate that the proposed end-to-end structure can outperform much better than the cascaded baselines.