Leveraging NLP for Building Efficient Information Retrieval Systems: A Performance Analysis
Chen Chung Tung, Zhi Tan, Zhen Yew Chook, Chi Wee Tan, Khai Yim Lim · 2024
The advancement of technology has highlighted the inefficiencies of traditional methods for extracting information from written materials such as PDFs. This paper explores an interactive chat application designed to enhance interaction with PDF content using Large Language Models (LLMs). The application aims to streamline the extraction process by allowing users to query PDF content efficiently. This project evaluates various models, including RoBERTa, TinyRoBERTa, and MDeBERTa, for their performance in question-and-answer tasks, and assesses their multilingual capability. The research questions investigate the effectiveness and language support of these models, with hypotheses predicting that RoBERTa will excel in accuracy, MDeBERTa will perform well with Chinese queries. Results reveal that RoBERTa achieved the highest accuracy for English queries, MDeBERTa performed best with Chinese queries. This project demonstrates the potential of these technologies to improve information extraction and user interaction with documents. Future work will focus on addressing the system’s limitations with mixed languages and non-text inputs to enhance its overall functionality.