Research on the Construction of Multimodal Datasets for Digital Libraries
Yi Zeng, Juxiang Zhou, Tianwei Xu · 2023
In order to apply cross-modal retrieval to digital library retrieval services, this paper firstly collects image and text data in digital libraries by crawlers and filters out the unqualified data; then adds text descriptions for images without text descriptions; and finally labels the images using annotation tools. A multimodal dataset containing 4400 image-text pairs is constructed and experimentally validated on several cross-modal retrieval methods. The experiments show that the dataset constructed in this paper is suitable for cross-modal retrieval research.