Jiajun Bao
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
2026/02/13 by Xiangyi Li, Yimin Liu, Wenbo Chen +75 · 27 voices · 53 citations
Computer Science · #cs.AI
- ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
2026/04/06 by Xiangyi Li, Kyoung Whan Choe, Yimin Liu +12 · 1 voice · 2 citations
Computer Science · #cs.AI