Changzhi Zhou
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
2026/01/17 by Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82 · 3 voices · 12 citations
Computer Science · #cs.SE #cs.AI
- ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
2025/07/07 by Chenchen Zhang, Yuhang Li, Zhang, Chenchen +36 · 8 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Software Engineering (cs.SE)
- A Comprehensive Evaluation of Large Language Models on Aspect-Based Sentiment Analysis
2024/12/03 by Changzhi Zhou, Dandan Song, Zhou, Changzhi +15 · 2 citations
Computer Science · #Advanced Text Analysis Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Sentiment Analysis and Opinion Mining #Text and Document Classification Technologies