vix.ing · top · new · best · stats · spec

Mateusz Dziemian

  1. AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
    2024/10/11 by Maksym Andriushchenko, Alexandra Souly, Andriushchenko, Maksym +25 · 56 citations
    Computer Science · #Blockchain Technology Applications and Security
  2. Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
    2025/07/28 by Andy Zou, Zou, Andy, Maxwell Lin +31 · 5 voices · 8 citations
    #cs.AI #cs.CL #cs.CY
  3. Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
    2025/04/10 by Simon Lermen, Mateusz Dziemian, Lermen, Simon +3 · 1 voice · 2 citations
    #cs.AI #cs.CL