Neev Parikh
- Measuring AI Ability to Complete Long Software Tasks
2025/03/18 by Thomas Kwa, Kwa, Thomas, Ben West +48 · 26 voices · 47 citations
#cs.AI #cs.LG
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
2024/11/22 by Hjalmar Wijk, Wijk, Hjalmar, Tao Lin +47 · 3 voices · 25 citations
Computer Science · Medicine · Social Sciences · #Artificial Intelligence in Healthcare and Education #Ethics and Social Impacts of AI #Explainable Artificial Intelligence (XAI) #cs.AI #cs.LG
- Learning Markov State Abstractions for Deep Reinforcement Learning
2021/06/08 by Cameron Allen, Allen, Cameron, Neev Parikh +5 · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics