Mateusz Dziemian
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
2024/10/11 by Maksym Andriushchenko, Alexandra Souly, Andriushchenko, Maksym +25 · 56 citations
Computer Science · #Blockchain Technology Applications and Security
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
2025/07/28 by Andy Zou, Zou, Andy, Maxwell Lin +31 · 5 voices · 8 citations
#cs.AI #cs.CL #cs.CY
- Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
2025/04/10 by Simon Lermen, Mateusz Dziemian, Lermen, Simon +3 · 1 voice · 2 citations
#cs.AI #cs.CL