- ORCA-bench: How Ready Are Language Model Agents for Oncall?
2026/07/30 by Albert Gong, Kyuseong Choi, Abhineet Agarwal +5 · 3 voices
Computer Science · #cs.AI #cs.CL #cs.SE
- Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution
2026/07/30 by I. Kennedy, T. Kennedy · 2 voices
Computer Science · #cs.CL
- Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B
2026/07/30 by Iliya Mirzaei · 1 voice
Computer Science · #cs.AI #cs.CL #cs.LG
- AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
2026/07/30 by Xiangning Lin, Shenzhe Zhu, Shu Yang +23 · 1 voice
Computer Science · #cs.AI #cs.CL #cs.CY #cs.HC
- Inducing language models to assert their own consciousness restores human beliefs and values
2026/07/30 by Junsol Kim, Winnie Street, Roberta Rocca +4 · 2 voices
Computer Science · #cs.CL
- AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
2026/07/30 by Bing Yan, Gregory Wolfe, Stefano Martiniani +1 · 1 voice
Computer Science · #cs.AI #cs.CL #cs.IR #cs.LG
- Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
2026/07/30 by Junlin Yang, Che Jiang, Yu Fu +21 · 1 voice
Computer Science · #cs.CL
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
2026/07/30 by Qiushi Sun, Kanzhi Cheng, Yian Wang +20 · 1 voice
Computer Science · #cs.AI #cs.CL #cs.CV
- Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
2026/07/29 by Alexi Gladstone, Heng Ji, Yilun Du · 1 voice
Computer Science · #cs.LG #cs.AI #cs.CL #cs.CV
- HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
2026/07/28 by Liudas Panavas, Sebastian Minus, Bradley Monton +4 · 3 voices
Computer Science · #cs.AI #cs.CL
- Metis: Memory Foundation Model
2026/07/29 by Zeyu Zhang, Ziliang Guo, Yihang Sun +14 · 2 voices
Computer Science · #cs.CL #cs.LG
- Hearsay: Vision-Language Medical Diagnoses Without an Image
2026/07/29 by Siddharth Vohra · 1 voice
Computer Science · #cs.CV #cs.AI #cs.CL #cs.CY
- Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors
2026/07/28 by Jianfei Ma, Zhaoxin Feng, Emmanuele Chersoni +1 · 1 voice
Computer Science · #cs.AI #cs.CL
- A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series
2026/07/28 by Frank Nie, Ethan B Liu, Yuan Zhu +2 · 1 citation
#cs.AI #cs.CL
- PILA: Plug-and-Play Insertion for LLM-native Advertising
2026/07/28 by Zhaowei Zhang, Yuhan Fu, Yihang Zhang +6 · 1 citation
#cs.CL
- Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
2026/07/27 by Tapan Parikh · 1 voice
#cs.CL #cs.AI
- Kimi K3: Open Frontier Intelligence
2026/07/27 by Kimi Team, Tongtong Bai, Yifan Bai +398 · 1 voice
#cs.CL #cs.LG
- A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever
2026/07/26 by Sietse Schelpe · 2 voices
#cs.CL #cs.AI #cs.IR #cs.LG #cs.PF
- Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings
2026/07/24 by Quentin Spencer · 2 voices
#cs.CL
- The Boundaries of Automation: A Theory of Persistent Human Participation
2026/07/23 by Fares Fourati, Hinrich Schütze, Eyke Hüllermeier +1 · 1 voice
#cs.AI #cs.CL #cs.ET #cs.LG #cs.MA
- Error Certificates for KV-Cache Eviction via Randomized Design
2026/07/23 by Peng Xie · 1 voice
#cs.LG #cs.AI #cs.CL
- When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs
2026/07/23 by Anna Mosolova, Djamé Seddah · 1 voice
#cs.CL
- Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks
2026/07/23 by Hiroki Tamba · 1 voice
#cs.LG #cs.CL
- Agentic Evaluation of Copyright Law Compliance
2026/07/23 by Zheng Hui, Doni Bloomfield, Noam Kolt · 1 citation
#cs.CL #cs.CY
- Generative AI floods and dilutes the market for books
2026/07/22 by Tuhin Chakrabarty, Xinyue Liu, Jane C. Ginsburg +1 · 3 voices
#cs.CL #cs.AI #cs.CY
- Notes to Self: Can LLMs Benefit from Experiential Abstractions?
2026/07/22 by Chang Liu, Xinyu Li, Artur Dubrawski · 1 voice
#cs.CL
- Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking
2026/07/22 by Kailin Jiang, Lei Liu, Jian Xi +7 · 1 citation
#cs.CL
- surprisal is Not a Theory
2026/07/22 by Andrés Buxó-Lugo, Aniello De Santo, Morgan Grobol +2 · 1 voice
#cs.CL
- RAGAL: A Frugal, Fully Local Retrieval-Augmented Assistant for Technical Support at a Government Agency
2026/07/21 by Dan Musetoiu · 1 voice
#cs.IR #cs.CL
- Measuring Reward-Seeking via Contrastive Belief Updates
2026/07/21 by Axel Højmark, Jérémy Scheurer, Evgenia Nitishinskaya +5 · 1 voice · 1 citation
#cs.AI #cs.CL #cs.LG
more