Hacking LLMs: A Technical Analysis of Security Vulnerabilities and Defense Mechanisms
Gaurav Raj, Hamzah, Nikhil Raj, Nikhil Ranjan · 2025
Large Language Models (LLMs) such as GPT-4 and Google’s Gemini have revolutionized the landscape of artificial intelligence, enabling sophisticated natural language processing capabilities across diverse domains. However, the rapid adoption of these models has unveiled a spectrum of security vulnerabilities exploitable by adversaries [1], [2]. This paper presents a comprehensive analysis of various attack vectors targeting LLMs, including prompt injection, data poisoning, model inversion, and side-channel attacks. We examine real-world incidents and vulnerabilities discovered in widely-used LLMs like ChatGPT and Gemini [3], highlighting the intricate mechanisms through which these attacks compromise model integrity and user trust. Furthermore, we explore existing defense mechanisms and best practices, emphasizing runtime protection systems, adversarial training approaches, and robust access control systems [4], [5]. Building upon these foundations, we introduce a novel Multi-Layer Defense Framework with Dynamic Trust Scoring, designed to enhance the security and robustness of LLM deployments through dynamic evaluation and adaptive response systems. By dissecting recent advancements in LLM security research and proposing innovative defense strategies, this paper aims to guide the development of resilient defenses against the evolving landscape of AI-driven threats.