Designing Prompts and Creating Cleaned Scientific Text for Retrieval Augmented Generation for More Precise Responses from Generative Large Language Models

Róbert Lakatos, Eszter Klára Urbán, Zoltán János Szabó, János Pozsga, Eszter Csernai, András Hajdú · 2024

This paper presents a comprehensive methodology for extracting and processing data from the scientific literature to improve the performance of generative language models in the case of the application of Retrieval Augmented Generation. We show how a knowledge-based system can be created to extract information from scientific literature using generative language models. The methodology involves a two-phase approach, utilizing the GROBID PDF processing system for initial data extraction, followed by refinement through a custom text-cleaning module. The processed data is formatted into JSON for integration into a semantic search engine, facilitating efficient searching and retrieval. Additionally, the setup of prompts and management of generative language models are meticulously detailed to optimize response quality. Evaluation of response performance using metrics such as BLEU, ROUGE, METEOR, and cosine similarity demonstrates the efficiency of the proposed methodology. As a result, the efficiency of generative language models using data embedded and cleaned in our knowledge-based system improves by 9 % in terms of cosine similarity and by 27 %, 6%, and 2% in the case of BLEU, ROUGE, METEOR scores, in contrast to the direct usage of the data extracted by GROBID. Overall, this work showcases the methodology's effectiveness in improving the quality of responses generated by language models and lays the groundwork for further advancements in natural language processing and semantic search systems within the scientific literature domain.

Read the paper · More papers on PaperTik