2024/12/01 by Chen-Wei Chang, Chang, Chen-Wei, Shailik Sarkar +16 · 2 citations
Computer Science · #Advanced Malware Detection Techniques #Artificial Intelligence (cs.AI) #Computers and Society (cs.CY) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Network Security and Intrusion Detection #Social and Information Networks (cs.SI) #Spam and Phishing Detection
paper · pdf · doi:10.48550/arxiv.2412.00621
openalex publication_date 2024/12/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Can we trust Large Language Models (LLMs) to accurately predict scam? This paper investigates the vulnerabilities of LLMs when facing adversarial scam messages for the task of scam detection. We addressed this issue by creating a comprehensive dataset with fine-grained labels of scam messages, including both original and adversarial scam messages. The dataset extended traditional binary classes for the scam detection task into more nuanced scam types. Our analysis showed how adversarial examples took advantage of vulnerabilities of a LLM, leading to high misclassification rate. We evaluated the performance of LLMs on these adversarial scam messages and proposed strategies to improve their robustness.