Carlo Leonardo Attubato
- Probing and Steering Evaluation Awareness of Language Models
2025/07/02 by Jord Nguyen, Khiem Hoang, Nguyen, Jord +5 · 6 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences
- Benchmarking Deception Probes via Black-to-White Performance Boosts
2025/07/16 by Avi Parrack, Parrack, Avi, Carlo Leonardo Attubato +3 · 4 citations
Computer Science · #Advanced Malware Detection Techniques #Adversarial Robustness in Machine Learning #Network Security and Intrusion Detection