SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
2026/02/13 by Xiangyi Li, Yimin Liu, Wenbo Chen +75 · 27 voices · 24 citations
#cs.AI
paper · pdf
Abstract
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to measure whether they actually help. We present SkillsBench, a benchmark whose current inventory contains 87 tasks across 8 domains paired with curated Skills and deterministic verifiers. Our latest aggregate evaluation runs the 87-task benchmark under matched no-Skills and curated-Skills conditions for 18 model-harness configurations. Curated Skills raise the average pass rate from 33.9% to 50.5% (+16.6 percentage points; 25.5% normalized gain), with configuration-level gains ranging from +4.1 to +25.7 pp. Focused Skills with at most three modules outperform larger or exhaustive bundles, and smaller models with Skills can match larger models without them. SkillsBench establishes paired evaluation as the foundation for rigorous measurement of Skill efficacy on agentic, expertise-heavy work.
Citations
Cited by
Discussions
- SkillsBench: Benchmarking how well agent skills work across diverse tasks [hn, 364 points, 171 comments]
- arxiv.org/abs/2602.12670 This is why hexdocs.pm/usage_rules exists. Package authors need to be providing the context that will guide your agent to successfully use their package. As their library chan [bsky, 14 points, 2 comments]
- https://bsky.app/profile/ruoshuiresearch.bsky.social/post/3mg27twhxfc2n [bsky, 3 points, 1 comments]
- https://bsky.app/profile/tommis.fi/post/3mfixwvfyzc2a [bsky, 2 points, 0 comments]
- SkillsBench: Benchmarking How Well Agent #Skills Work Across Diverse Tasks https://arxiv.org/abs/26... #LLM #ClaudeCode [bsky, 1 points, 0 comments]
- Study: Self-generated Agent Skills are useless [bsky, 1 points, 0 comments]
- Skills benchmark and evals, main contributions are 1. skills benchmark 2. self-generated Skills are not so great arxiv.org/pdf/2602.12670 [bsky, 1 points, 1 comments]
- Study: Self-generated Agent Skills are useless #HackerNews https://arxiv.org/abs/2602.12670 [bsky, 0 points, 0 comments]
- Study: Self-generated Agent Skills are useless https://arxiv.org/abs/2602.12670 https://news.ycombinator.com/item?id=47040430 [bsky, 0 points, 0 comments]
- Study: Self-generated Agent Skills are useless https://arxiv.org/abs/2602.12670 https://news.ycombinator.com/item?id=47040430 [bsky, 0 points, 0 comments]
- SkillsBench: Benchmarking how well agent skills work across diverse tasks https://arxiv.org/abs/2602.12670 comments #arxiv.org [bsky, 0 points, 0 comments]
- Study: Self-generated Agent Skills are useless https://arxiv.org/abs/2602.12670 [bsky, 0 points, 0 comments]
- Similar take for skills arxiv.org/pdf/2602.12670 [bsky, 0 points, 0 comments]
- 📰 Study: Self-generated Agent Skills are useless 🔗 https://arxiv.org/abs/2602.12670 💬 Discuss on HN [bsky, 0 points, 0 comments]
- Paper: arxiv.org/pdf/2602.12670 [bsky, 0 points, 0 comments]
- ⚡ Hackernews Top story: Study: Self-generated Agent Skills are useless [bsky, 0 points, 0 comments]
- Study: Self-generated Agent Skills are useless https://arxiv.org/abs/2602.12670 (http://news.ycombinator.com/item?id=47040430) [bsky, 0 points, 0 comments]
- Study: Self-generated Agent Skills are useless https://arxiv.org/abs/2602.12670 (http://news.ycombinator.com/item?id=47040430) [bsky, 0 points, 0 comments]
- https://arxiv.org/abs/2602.12670 SkillsBenchは、AIエージェントの能力を測る新しいベンチマークです。 多様なタスクにおけるエージェントのスキル性能を評価します。 AIエージェントの汎用性を理解する上で重要な指標を提供します。 [bsky, 0 points, 0 comments]
- Study: Self-generated Agent Skills are useless view on hacker news [bsky, 0 points, 0 comments]
- SkillsBench: Benchmarking how well agent skills work across diverse tasks https:// arxiv.org/abs/2602.12670 # arxiv [mastodon, 0 points, 0 comments]
- SkillsBench: Benchmarking how well agent skills work across diverse tasks View Article | Join the HN Conversation Summary of HN discussion 🧵👇 [bsky, 0 points, 1 comments]
- Study: Self-generated Agent Skills are useless https://arxiv.org/abs/2602.12670 (https://news.ycombinator.com/item?id=47040430) [bsky, 0 points, 0 comments]
- Study: Self-generated Agent Skills are useless https://arxiv.org/abs/2602.12670 (https://news.ycombinator.com/item?id=47040430) [bsky, 0 points, 0 comments]
- Study: Self-generated Agent Skills are useless https://arxiv.org/abs/2602.12670 (https://news.ycombinator.com/item?id=47040430) [bsky, 0 points, 0 comments]
- https://bsky.app/profile/buzzing.cc.web.brid.gy/post/3mezbkpy4bil2 [bsky, 0 points, 0 comments]
- 1/2 Estudio: las habilidades auto-generadas por agentes no sirven Un paper de arXiv analiza si los agentes que se crean sus propias habilidades mejoran realmente el rendimiento. La discusión en Hacker [bsky, 0 points, 1 comments]
Related