- ORCA-bench: How Ready Are Language Model Agents for Oncall?
2026/07/30 by Albert Gong, Kyuseong Choi, Abhineet Agarwal +5 · 3 voices
Computer Science · #cs.AI #cs.CL #cs.SE
- MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems
2026/07/30 by Mao-xun Huang, Jerry Wang, Yi-Cheng Lai +3 · 1 voice
Computer Science · #cs.AI
- Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B
2026/07/30 by Iliya Mirzaei · 1 voice
Computer Science · #cs.AI #cs.CL #cs.LG
- AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
2026/07/30 by Xiangning Lin, Shenzhe Zhu, Shu Yang +23 · 1 voice
Computer Science · #cs.AI #cs.CL #cs.CY #cs.HC
- AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
2026/07/30 by Bing Yan, Gregory Wolfe, Stefano Martiniani +1 · 1 voice
Computer Science · #cs.AI #cs.CL #cs.IR #cs.LG
- Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation
2026/07/30 by Marco Alecci, Francesco Marchiori, Iyiola Emmanuel Olatunji +2 · 1 voice
Computer Science · #cs.AI #cs.CR #cs.SE
- Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
2026/07/30 by Hanzhang Zhou, Panrong Tong, Xu Zhang +13 · 1 voice
Computer Science · #cs.AI #cs.CV
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
2026/07/30 by Qiushi Sun, Kanzhi Cheng, Yian Wang +20 · 1 voice
Computer Science · #cs.AI #cs.CL #cs.CV
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
2026/07/29 by Alexi Gladstone, Heng Ji, Yilun Du · 1 voice
Computer Science · #cs.LG #cs.AI #cs.CL #cs.CV
- HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
2026/07/28 by Liudas Panavas, Sebastian Minus, Bradley Monton +4 · 3 voices
Computer Science · #cs.AI #cs.CL
- Can AI agents conduct open-ended AI research? Early evidence from two case studies
2026/07/29 by Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21 · 2 voices
Computer Science · #cs.AI #cs.CY #cs.LG
- Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction
2026/07/29 by Xin Xu, Chengrui Wu, Jiayu Lu +3 · 2 voices
Computer Science · #cs.AI #cs.CR #cs.GT
- A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
2026/07/29 by Wenhao Yang, Runzhi He, Minghui Zhou · 1 voice
Computer Science · #cs.SE #cs.AI
- Hearsay: Vision-Language Medical Diagnoses Without an Image
2026/07/29 by Siddharth Vohra · 1 voice
Computer Science · #cs.CV #cs.AI #cs.CL #cs.CY
- Human diversity fuels collective creativity that large language models cannot simulate or sustain
2026/07/29 by Mengchen Dong, Hiromu Yakura · 1 voice
Computer Science · #cs.HC #cs.AI #cs.CY
- Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
2026/07/28 by Jiabao Ji, Yujian Liu, Li An +4 · 1 voice
#cs.AI
- Visual prompt engineering for video models
2026/07/28 by Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7 · 1 voice
#cs.CV #cs.AI
- Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
2026/07/28 by Stefan Krsteski, Charlotte Meyer, Guillaume Allegre +2 · 1 voice
#cs.AI #cs.DB
- Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors
2026/07/28 by Jianfei Ma, Zhaoxin Feng, Emmanuele Chersoni +1 · 1 voice
Computer Science · #cs.AI #cs.CL
- Specula: Scaling formal specifications for autonomous model checking of system code
2026/07/28 by Qian Cheng, Saad Mohammad Rafid Pial, Ruize Tang +6 · 1 voice
#cs.SE #cs.AI #cs.DC #cs.OS
- At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference
2026/07/28 by Bowen Wang, Chi Zhang, Diyou Shen +3 · 1 voice
#cs.AR #cs.AI
- A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series
2026/07/28 by Frank Nie, Ethan B Liu, Yuan Zhu +2 · 1 citation
#cs.AI #cs.CL
- From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
2026/07/27 by Junlin Liu, Jiangwang Chen, Zixin Song +7 · 2 voices · 1 citation
#cs.AI
- Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
2026/07/27 by Tapan Parikh · 1 voice
#cs.CL #cs.AI
- Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents
2026/07/27 by Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov +2 · 1 voice
#cs.CR #cs.AI
- The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing
2026/07/27 by Stefan Scholze, Johannes Partzsch, Sebastian Höppner +27 · 1 voice
#cs.ET #cs.AI #cs.AR #cs.DC
- A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever
2026/07/26 by Sietse Schelpe · 2 voices
#cs.CL #cs.AI #cs.IR #cs.LG #cs.PF
- An Exact Counterexample to Carlson's Associated-Prime Depth Conjecture from a Group of Order 128
2026/07/26 by Xinan Dai, Wenhao Deng, Yingdong Shi +2 · 1 citation
#math.GR #cs.AI
- Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
2026/07/24 by Jiaqi Shao, Hanck Chen, Wei Zhang +2 · 1 voice
#cs.AI
- Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
2026/07/23 by Gaurav Dadhich · 2 voices
#cs.AI #cs.IR
more