MuseumQA: A Fine-Grained Question Answering Dataset for Museums and Artifacts

Fan Jin, Qingling Chang, Zhongwen Xu · 2023

In this paper, we present a fine-grained museum artifact question-answering (QA) dataset, which serves as the cornerstone for developing museum question-answering systems. Creating these systems is essential for the advancement of museums and can enhance the visitor experience. Nevertheless, research reveals the current absence of domestically available datasets for museum artifacts in China. To ensure data authenticity and validity, we meticulously collected and screened 3,416 raw data entries from the official websites of provincial museums across China. Using these raw data, we annotated annotatable QA pair information to create the final QA dataset. Initially, a small batch of QA pairs was generated with the assistance of ChatGPT. Subsequently, the remaining QA pairs were annotated using an enhanced QA generation model, yielding 23,149 QA pairs. To mitigate overfitting due to dataset-model size disparities, a noise factor was incorporated into the enhanced generation model. Additionally, a Chinese grammar correction module was integrated to enhance the accuracy of the generated statements. Ultimately, the model achieved optimal performance, and the dataset demonstrated the highest semantic relevance.

Read the paper · More papers on PaperTik