Designing a structured approach for consistency verification in AI systems for heartbeat disorder-related conversations
Muhammad Badruddin Khan, Abdul Khader Jilani Saudagar · Journal of Statistics and Management Systems · 2024
Large language models (LLMs) have shown mind-boggling question-answering capabilities. Most of their responses make one feel that they understand and remember context and are able to retrieve relevant content to generate almost accurate, human-like responses. While LLMs seem to mimic human intelligence, their responses can sometimes be inconsistent. Although LLMs are getting acceptance in healthcare for tasks like diagnosis support or patient communication, their inaccurate, inconsistent, or misleading information can be dangerous and can lead to harmful medical decisions if not properly validated by experts. Due to these limitations, LLM-powered full automation is still a dream. This paper presents a structured approach that can be used as basis to check LLMs level of “understanding” of medical knowledge and their “inferential capabilities” by evaluation of their responses to the word problems related to human heartbeat. The crafted word problems designed in the light of the proposed approach can be very helpful in identifying inherent weaknesses stemming from their fundamental nature. The study can be highly beneficial for healthcare professionals in determining the appropriate level of adoption of modern artificial intelligence (AI) technologies to identify heartbeat related disorders like Arrhythmia.