Conversational open-domain question answering for resource-constrained languages
Emrah Budur, Tunga Güngör · TURKISH JOURNAL OF ELECTRICAL ENGINEERING & COMPUTER SCIENCES · 2025
The growing interest in Conversational AI has led to the development of Conversational OpenQA systems as a crucial step for meeting users' information needs in real world scenarios. Conversational OpenQA systems enhance standard OpenQA performance by leveraging conversation history of the users. However, building effective Conversational OpenQA systems requires large-scale Conversational OpenQA datasets, often limited to the English language, hindering progress in low-resource languages. We present a robust Conversational OpenQA system enhanced by conversational context, designed for languages with limited resources and exemplified in our case study for Turkish. To address data limitations in a cost-effective way, we repurpose existing datasets like SQuAD-TR and XQuAD-TR, treating them as if they were constructed within a conversational context. Our findings indicate that incorporating conversation signals in the retriever models results in up to an absolute increase of 18.82% in Success@1 for retrievers. This improvement extends to the reader models enhanced by the conversational context, narrowing the gap in EM/F1 scores up to 4.12% / 4.43%, respectively, compared to Standard QA readers.