Michal Shmueli-Scheuer
- AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
2026/06/11 by Xiaoyuan Liu, Jianhong Tu, Yuqi Chen +26 · 1 voice · 1 citation
Computer Science · #cs.AI #cs.LG
- Robustness as an Emergent Property of Task Performance
2026/02/03 by Shir Ashury-Tahan, Ariel Gera, Elron Bandel +2 · 1 voice
Computer Science · #cs.LG #cs.AI #cs.CL
- Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data
2026/07/09 by Ofir Arviv, Kristjan Greenewald, Yotam Perlitz +3 · 1 voice
Computer Science · #cs.LG
- Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
2026/04/14 by Eliya Habba, Itay Itzhak, Asaf Yehudai +5 · 1 voice
Computer Science · #cs.CL