LocatingGPT: A multi-modal document retrieval method based on retrieval-augmented generation
Zilong Chen, Peng Zhang, Mingyang Xu, Xingjian Gao, Dan Liu, Xuebing Luo, Rencheng Huang, Tao Zhang · 2024
With the rapid development of information technology, document information retrieval systems are faced with challenges of data diversification and increasing complexity. Conventional document retrieval approaches show intrinsic limitations when addressing multi-modal data (such as text, images, and audio), cross-language queries, and complex queries, failing to meet the needs of modern information environments. To tackle these issues, we propose LocatingGPT, a multi-modal document retrieval system based on the Retrieval-Augmented Generation framework. This system amalgamates the Donut and Whisper models, facilitating the efficient retrieval and localization of diverse data types, including text, images, and audio. It particularly emphasizes the adept handling of cross-language queries, thereby supporting document parsing and information retrieval in multilingual environments. Additionally, by integrating a deep parsing module, LocatingGPT achieves remarkable proficiency in indexing documents of low similarity and in the nuanced reasoning of complex queries, thus significantly enhancing the relevance and accuracy of the retrieval results. Experimental results indicate that LocatingGPT demonstrates efficiency and reliability in multi-modal document retrieval and cross-language queries. When addressing ambiguous queries and complex knowledge reasoning tasks, this method enhances the performance of Retrieval Question Answering (ReQA) tasks without appreciably increasing computational demands, showcasing its robustness.