2025/02/03 by Meenatchi Sundaram Muthu Selva Annamalai, Annamalai, Meenatchi Sundaram Muthu Selva, Igor Bilogrevic +3 · 1 voice · 1 citation
Computer Science · Social Sciences · #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #Hate Speech and Cyberbullying Detection #Human-Computer Interaction (cs.HC) #Privacy, Security, and Data Protection #User Authentication and Security Systems #cs.CR #cs.HC
paper · pdf · doi:10.48550/arxiv.2502.01608
openalex publication_date 2025/02/03 · arxiv published 2025/02/03 · arxiv updated 2025/02/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Browser fingerprinting is a pervasive online tracking technique used increasingly often for profiling and targeted advertising. Prior research on the prevalence of fingerprinting heavily relied on automated web crawls, which inherently struggle to replicate the nuances of human-computer interactions. This raises concerns about the accuracy of current understandings of real-world fingerprinting deployments. As a result, this paper presents a user study involving 30 participants over 10 weeks, capturing telemetry data from real browsing sessions across 3,000 top-ranked websites. Our evaluation reveals that automated crawls miss almost half (45%) of the fingerprinting websites encountered by real users. This discrepancy mainly stems from the crawlers' inability to access authentication-protected pages, circumvent bot detection, and trigger fingerprinting scripts activated by specific user interactions. We also identify potential new fingerprinting vectors present in real user data but absent from automated crawls. Finally, we evaluate the effectiveness of federated learning for training browser fingerprinting detection models on real user data, yielding improved performance than models trained solely on automated crawl data.