2018/04/19 by Dominic Seyler, Seyler, Dominic, Lunan Li +3
Computer Science · #Computation and Language (cs.CL) #Cryptography and Security (cs.CR) #Cybercrime and Law Enforcement Studies #FOS: Computer and information sciences #Network Security and Intrusion Detection #Social and Information Networks (cs.SI) #Spam and Phishing Detection #cs.CL #cs.CR #cs.SI
paper · pdf · doi:10.48550/arxiv.1804.07247
openalex publication_date 2018/04/19 · arxiv created 2020/10/23 · arxiv updated 2020/10/26 · openalex created_date 2022/10/02 · openalex updated_date 2026/07/28
Compromised accounts on social networks are regular user accounts that have been taken over by an entity with malicious intent. Since the adversary exploits the already established trust of a compromised account, it is crucial to detect these accounts to limit the damage they can cause. We propose a novel general framework for semantic analysis of text messages coming out from an account to detect compromised accounts. Our framework is built on the observation that normal users will use language that is measurably different from the language that an adversary would use when the account is compromised. We propose to use the difference of language models of users and adversaries to define novel interpretable semantic features for measuring semantic incoherence in a message stream. We study the effectiveness of the proposed semantic features using a Twitter data set. Evaluation results show that the proposed framework is effective for discovering compromised accounts on social networks and a KL-divergence-based language model feature works best.