TiconvQA: A Tibetan Conversational Dataset for Text Comprehension
Pengmao Cairang, Dawa Cairen, Yuan Sun · Data Intelligence · 2025
Conversational machine comprehension is a research area in conversational artificial intel-ligence, aiming to enable machines to understand text and engage in multi-turn dialogues to answer questions. In recent years, the emergence of multi-turn question answering datasets and the development of pre-trained language models have attracted widespread attention in the field of multi-turn question answering, further driving the advancement of conversational machine comprehension. This paper introduces TiconvQA (Tibetan Conversational Question Answering), a new Tibetan dataset for research on Tibetan conversational question answering tasks. The dataset contains 10,000 questions with answers, which come from 977 text paragraphs in three domains: people, geography, and news. The questions are conversational in nature, the answers are free-form text, and evidence text is provided within the article. The release of this dataset fills a gap in the field of conversational machine comprehension for Tibetan, and has significant implications for promoting the development of Tibetan information processing.