An improved zero-resource black-box framework for detecting hallucinations in 100 billion+ parameter generative pre-trained large language models
Jokie Blaise, Fatima Umar Zambuk, Badamasi Imam Ya’u · 2024
Large Language Models (LLMs) have garnered widespread attention due to their impressive performance, particularly in text generation. There is a strong inclination to apply these models across various fields. However, these models are not without their vulnerabilities, with hallucination being a prominent concern. Hallucination occurs when the model generates outputs that are either non-factual or lack grounding in the source input, while maintaining fluency and grammatical correctness. This drawback diminishes the trustworthiness of these models, rendering them unsuitable for deployment in sensitive fields and susceptible to misuse for disinformation. As the size of these models continues to grow, and the competition to unveil the next breakthrough model persists, developers often release them in a closed format, keeping their internal workings opaque. The lack of transparency underscores the importance of black box approaches to hallucination detection. In response to recent developments in zero-resource, black-box hallucination detection approaches, this work aims to enhance these methods, contributing valuable tools to address hallucination issues in large language models. We introduce a zero-resource, black-box framework named HalluciCheck, which operates based on an ensemble of similarity metric scores. We find that this framework has an improved performance in detecting factual sentences while its performance in detecting false sentences is comparable to previous methods. This framework serves as an initial step toward mitigating the hallucination problem.