vix.ing · top · new · best · stats · spec

Zhongyuan Peng

  1. SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
    2025/02/20 by M-A-P Team, P Team, Xinrun Du +179 · 71 citations
    Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences
  2. RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
    2023/10/01 by Zekun Moore Wang, Zhongyuan Peng, Wang, Zekun Moore +29 · 27 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  3. A Comparative Study on Reasoning Patterns of OpenAI's o1 Model
    2024/10/17 by Siwei Wu, Zhongyuan Peng, Wu, Siwei +31 · 8 citations
    Business, Management and Accounting · Computer Science · #Advanced Database Systems and Queries #Business Process Modeling and Analysis #Computation and Language (cs.CL) #FOS: Computer and information sciences #Semantic Web and Ontologies
  4. FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
    2025/05/05 by Zhouliang Yu, Ruotian Peng, Yu, Zhouliang +23 · 10 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
  5. MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
    2024/10/15 by Pei Wang, Yanan Wu, Wang, Pei +27 · 3 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  6. CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
    2025/02/23 by Alexander Zhang, Zhang, Alexander, Jiaheng Liu +32 · 4 citations
    Computer Science · #Software Engineering Research #Software System Performance and Reliability #Topic Modeling
  7. ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders
    2026/07/23 by Zhongyuan Peng, Dan Huang, Chuyu Zhang +8
    #cs.AI