Jason Z Wang
- The Evaluation Blind Spot: A Stereological Theory of Benchmark Coverage for Large Language Models
2026/04/15 by Jason Z Wang · 1 voice
Computer Science · #cs.LG
- ERRORQUAKE: Heavy-Tailed Error Severity Distributions in Open-Weight Large Language Models
2026/04/15 by Jason Z Wang · 1 voice
Computer Science · #cs.LG
- MIRROR: A Hierarchical Benchmark for Metacognitive Calibration in Large Language Models
2026/04/15 by Jason Z Wang · 1 voice
Computer Science · #cs.AI #cs.LG