Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
2024/05/21 by Mohamed Amine Ferrag, Fatima Alwahedi, Ferrag, Mohamed Amine +11 · 28 citations
Computer Science · #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #Digital and Cyber Forensics #FOS: Computer and information sciences #Network Security and Intrusion Detection #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2405.12750
openalex publication_date 2024/05/21 · openalex created_date 2024/05/23 · openalex updated_date 2026/07/28
Abstract
This paper provides a comprehensive review of the future of cybersecurity through Generative AI and Large Language Models (LLMs). We explore LLM applications across various domains, including hardware design security, intrusion detection, software engineering, design verification, cyber threat intelligence, malware detection, and phishing detection. We present an overview of LLM evolution and its current state, focusing on advancements in models such as GPT-4, GPT-3.5, Mixtral-8x7B, BERT, Falcon2, and LLaMA. Our analysis extends to LLM vulnerabilities, such as prompt injection, insecure output handling, data poisoning, DDoS attacks, and adversarial instructions. We delve into mitigation strategies to protect these models, providing a comprehensive look at potential attack scenarios and prevention techniques. Furthermore, we evaluate the performance of 42 LLM models in cybersecurity knowledge and hardware security, highlighting their strengths and weaknesses. We thoroughly evaluate cybersecurity datasets for LLM training and testing, covering the lifecycle from data creation to usage and identifying gaps for future research. In addition, we review new strategies for leveraging LLMs, including techniques like Half-Quadratic Quantization (HQQ), Reinforcement Learning with Human Feedback (RLHF), Direct Preference Optimization (DPO), Quantized Low-Rank Adapters (QLoRA), and Retrieval-Augmented Generation (RAG). These insights aim to enhance real-time cybersecurity defenses and improve the sophistication of LLM applications in threat detection and response. Our paper provides a foundational understanding and strategic direction for integrating LLMs into future cybersecurity frameworks, emphasizing innovation and robust model deployment to safeguard against evolving cyber threats.
Cited by
- Improving Phishing Resilience with AI-Generated Training: Evidence on Prompting, Personalization, and Duration
- EAGER: Edge-Aligned LLM Defense for Robust, Efficient, and Accurate Cybersecurity Question Answering
- A Robust and Explainable Transformer-Based Framework for Phishing Email Detection
- Exploring Membership Inference Vulnerabilities in Clinical Large Language Models
- CLASP: Cost-Optimized LLM-based Agentic System for Phishing Detection
- Adversarial Defense in Cybersecurity: A Systematic Review of GANs for Threat Detection and Mitigation
- From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing
- AQUA-LLM: Evaluating Accuracy, Quantization, and Adversarial Robustness Trade-offs in LLMs for Cybersecurity Question Answering
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- SME-TEAM: Leveraging Trust and Ethics for Secure and Responsible Use of AI and LLMs in SMEs
- Neuro-Symbolic AI for Cybersecurity: State of the Art, Challenges, and Opportunities
- MultiFuzz: A Dense Retrieval-based Multi-Agent System for Network Protocol Fuzzing
- Can LLMs effectively provide game-theoretic-based scenarios for cybersecurity?
- An Agentic Flow for Finite State Machine Extraction using Prompt Chaining
- Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
- Federated Learning-Based Data Collaboration Method for Enhancing Edge Cloud AI System Security Using Large Language Models
- LLMs' Suitability for Network Security: A Case Study of STRIDE Threat Modeling
- SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis
- DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response
- Large Language Models in the IoT Ecosystem -- A Survey on Security Challenges and Applications
- Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks
- From Chasing Ghosts to Missed Attacks: Perspectives and Perceptions of SOC Practitioners on LLM Integration, Risks, and Readiness
- Towards Agentic Investigation of Security Alerts
- Digital Red Queen: Adversarial Program Evolution in Core War with LLMs
- The Role of Generative AI in Strengthening Secure Software Coding Practices: A Systematic Perspective
- Llama-3.1-FoundationAI-SecurityLLM-Base-8B Technical Report
- SoK: How Frontier AI Reshapes System-Level Security Risk Dynamics in Critical Infrastructure
- Digital Forensics in the Age of Large Language Models
Related