Zhongyuan Peng
- SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
2025/02/20 by M-A-P Team, P Team, Xinrun Du +179 · 71 citations
Social Sciences · #Artificial Intelligence in Law #Computation and Language (cs.CL) #FOS: Computer and information sciences
- RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
2023/10/01 by Zekun Moore Wang, Zhongyuan Peng, Wang, Zekun Moore +29 · 27 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
- A Comparative Study on Reasoning Patterns of OpenAI's o1 Model
2024/10/17 by Siwei Wu, Zhongyuan Peng, Wu, Siwei +31 · 8 citations
Business, Management and Accounting · Computer Science · #Advanced Database Systems and Queries #Business Process Modeling and Analysis #Computation and Language (cs.CL) #FOS: Computer and information sciences #Semantic Web and Ontologies
- FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
2025/05/05 by Zhouliang Yu, Ruotian Peng, Yu, Zhouliang +23 · 10 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
- MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
2024/10/15 by Pei Wang, Yanan Wu, Wang, Pei +27 · 3 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- CodeCriticBench: A Holistic Code Critique Benchmark for Large Language Models
2025/02/23 by Alexander Zhang, Zhang, Alexander, Jiaheng Liu +32 · 4 citations
Computer Science · #Software Engineering Research #Software System Performance and Reliability #Topic Modeling
- ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders
2026/07/23 by Zhongyuan Peng, Dan Huang, Chuyu Zhang +8
#cs.AI