vix.ing · top · new · best · stats · spec

Zhang, Chen Bo Calvin

  1. Humanity's Last Exam
    2025/01/24 by Long Phan, Alice Gatti, Phan, Long +2240 · 9 voices · 110 citations
    #cs.LG #cs.AI #cs.CL
  2. SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
    2025/09/21 by Xiang Deng, Deng, Xiang, Jeff Da +33 · 23 citations
    Business, Management and Accounting · Computer Science · #Business Process Modeling and Analysis #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Software Engineering (cs.SE)
  3. SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
    2025/06/17 by Kutasov, Jonathan, Sun, Yuqi, Colognese, Paul +9 · 11 citations
    Computer Science · Medicine · #Multimodal Machine Learning Applications #Artificial Intelligence in Healthcare and Education #Explainable Artificial Intelligence (XAI)
  4. ORSO: Accelerating Reward Design via Online Reward Selection and Policy Optimization
    2024/10/17 by Zhang, Chen Bo Calvin, Hong, Zhang-Wei, Pacchiano, Aldo +1 · 2 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Robotics (cs.RO)
  5. ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
    2025/11/10 by Sharma, Manasi, Zhang, Chen Bo Calvin, Bandi, Chaithanya +13 · 7 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
    2025/10/18 by Yu Ying Chiu, Chiu, Yu Ying, Michael S. Lee +30 · 2 citations
    Computer Science · Neuroscience · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Machine Learning (cs.LG) #Psychology of Moral and Emotional Judgment
  7. PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning
    2025/11/14 by Akyürek, Afra Feyza, Gosai, Advait, Zhang, Chen Bo Calvin +21 · 2 citations
    #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences