Between Truth and Hallucinations: Evaluation of the Performance of Large Language Model-Based AI Plugins in Website Quality Analysis

Karol Król · Applied Sciences · 2025

Although large language models (LLMs) like the Generative Pre-trained Transformer (GPT) are growing increasingly popular, much remains to learn about their potential for website quality auditing. The article evaluates the performance of LLM AI plugins (GPT models) in website and web application auditing. The author built and tested two original ChatGPT-4o Plus (OpenAI) plugins: Website Quality Auditor (WQA) and WebGIS Quality Auditor (WgisQA). Their performance was cautiously and carefully analysed and compared to traditional auditing tools. The results demonstrated the limitations of the AI plugins, including their propensity for false outcomes. The general conclusion is that using AI tools without considering their characteristics may lead to the propagation of AI hallucinations in audit reports. The study fills in the research gap with the results on the capabilities and limitations of AI plugins in the context of auditing. It also suggests further directions for improvement.

Read the paper · More papers on PaperTik