Enhancing Predictive Accuracy and Resource Utilization in Disease Prediction: A Comparative Study of Small Versus Large Language Models
V. Smrithi Rekha, A. Mohanarathinam, G. Sajiv, Natarajan Meenakshisundaram, P Praganya, Senthil Kumar C · 2024
This project involves a comparative analysis of small language and large language models where it evaluates the models’ capabilities in deciphering complex relationships between diseases and their symptoms using the distilgpt2(SLM) and gpt2(LLM) models. We utilized a large dataset where both small and large language models were used to study diseases and their symptoms. The models were evaluated based on their computational efficiency, predictive performance and the resources required for the operation. The methodology includes loading datasets, tokenization and model setup, monitoring training and validation losses, hyperparameter tuning to optimize the models and eventually generating text. The study showed SLM’s proficiency in producing context-aware responses which showed an inference time of 0.008518 secs whereas LLM shows its strength in generating refined, comprehensive text with a larger inference time of 0.015379 secs. The outcome showed the ability of large language models to possess higher predictive accuracy and they demanded significantly higher computational resources as compared to small language models, which may not be advised to use in resource-constrained environment. In contrast, small language models have efficient resource usage despite their low accuracy make their application more pertinent where computational resources are limited. This study paves the way for understanding the blueprint for health informatics practitioner highlighting the importance of domain-specific training in enhancing predictive accuracy and resource utilization of language models. It also underscores the requisite for a new approach to deployment of language models in healthcare settings.