Elephant: LLM System for Accurate Recantations
Praise Chinedu-Eneh, Trung Thanh Nguyen · 2024
This paper investigates the crossroads between Large Language Models (LLMs), Information Retrieval (IR), and the growing issue of misinformation in this current age of the Internet. Large Language Models, including the GPT and LLaMA series (both implemented in this paper), have given us profound insight into how we interact on the Internet and in real life. They have also innovated the way we interact with information. However, these innovations also pose significant shortfalls, particularly in the domain of misinformation from the text. The main contribution of this paper lies in developing a proposed strategy to mitigate the risks of misinformation seen on the Internet and generated from LLMs with a focus on public individuals by recanting statements they have made and the retrieval of said statements. We propose a multi-faceted approach that includes utilizing GPT3.5 and the open source LLaMA2 LLMs, finetuning data curation, and integrating accuracy mechanisms to ensure the most relevant and accurate information is retrieved. The efficacy of this methodology is measured using a cosine similarity metric. Considering that the recanting of this model must be at or as close to the original statement as possible, this metric is deemed most fitting. Findings later in this paper deemed a similarity recall of 90.42% on average with the GPT3.5 variant and 88.29% on average with the LLaMA2 variant, both in zero-shot examples, indicating the core semantic meanings were retrieved with variations on the format of illustration.