vix.ing · top · new · best · stats · spec

Jin, Tengjun

  1. Establishing Best Practices for Building Rigorous Agentic Benchmarks
    2025/07/03 by Zhu, Yuxuan, Jin, Tengjun, Pruksachatkun, Yada +22 · 5 voices · 12 citations
    #A.1 #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #I.2.m
  2. Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
    2025/10/13 by Sayash Kapoor, Kapoor, Sayash, Benedikt Stroebl +63 · 3 voices · 9 citations
    Computer Science · #Multi-Agent Systems and Negotiation
  3. ELT-Bench: An End-to-End Benchmark for Evaluating AI Agents on ELT Pipelines
    2025/04/07 by Yuxuan Zhu, Jin, Tengjun, Daniel Kang +2 · 5 citations
    Computer Science · Decision Sciences · #Advanced Database Systems and Queries #Artificial Intelligence (cs.AI) #Databases (cs.DB) #FOS: Computer and information sciences #Scientific Computing and Data Management #Semantic Web and Ontologies