vix.ing · top · new · best · stats · spec

Ryan Othniel Kearns

  1. Measuring what Matters: Construct Validity in Large Language Model Benchmarks
    2025/11/03 by Andrew M. Bean, Ryan Othniel Kearns, Bean, Andrew M. +83 · 2 voices · 14 citations
    Computer Science · Medicine · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #Computational and Text Analysis Methods #FOS: Computer and information sciences #Topic Modeling #cs.AI #cs.CL