vix.ing · top · new · best · stats · spec
  1. ORCA-bench: How Ready Are Language Model Agents for Oncall?
    2026/07/30 by Albert Gong, Kyuseong Choi, Abhineet Agarwal +5 · 3 voices
    Computer Science · #cs.AI #cs.CL #cs.SE
  2. Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution
    2026/07/30 by I. Kennedy, T. Kennedy · 2 voices
    Computer Science · #cs.CL
  3. Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B
    2026/07/30 by Iliya Mirzaei · 1 voice
    Computer Science · #cs.AI #cs.CL #cs.LG
  4. AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
    2026/07/30 by Xiangning Lin, Shenzhe Zhu, Shu Yang +23 · 1 voice
    Computer Science · #cs.AI #cs.CL #cs.CY #cs.HC
  5. Inducing language models to assert their own consciousness restores human beliefs and values
    2026/07/30 by Junsol Kim, Winnie Street, Roberta Rocca +4 · 2 voices
    Computer Science · #cs.CL
  6. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
    2026/07/30 by Bing Yan, Gregory Wolfe, Stefano Martiniani +1 · 1 voice
    Computer Science · #cs.AI #cs.CL #cs.IR #cs.LG
  7. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
    2026/07/30 by Junlin Yang, Che Jiang, Yu Fu +21 · 1 voice
    Computer Science · #cs.CL
  8. OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
    2026/07/30 by Qiushi Sun, Kanzhi Cheng, Yian Wang +20 · 1 voice
    Computer Science · #cs.AI #cs.CL #cs.CV
  9. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
    2026/07/29 by Alexi Gladstone, Heng Ji, Yilun Du · 1 voice
    Computer Science · #cs.LG #cs.AI #cs.CL #cs.CV
  10. HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
    2026/07/28 by Liudas Panavas, Sebastian Minus, Bradley Monton +4 · 3 voices
    Computer Science · #cs.AI #cs.CL
  11. Metis: Memory Foundation Model
    2026/07/29 by Zeyu Zhang, Ziliang Guo, Yihang Sun +14 · 2 voices
    Computer Science · #cs.CL #cs.LG
  12. Hearsay: Vision-Language Medical Diagnoses Without an Image
    2026/07/29 by Siddharth Vohra · 1 voice
    Computer Science · #cs.CV #cs.AI #cs.CL #cs.CY
  13. Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors
    2026/07/28 by Jianfei Ma, Zhaoxin Feng, Emmanuele Chersoni +1 · 1 voice
    Computer Science · #cs.AI #cs.CL
  14. A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series
    2026/07/28 by Frank Nie, Ethan B Liu, Yuan Zhu +2 · 1 citation
    #cs.AI #cs.CL
  15. PILA: Plug-and-Play Insertion for LLM-native Advertising
    2026/07/28 by Zhaowei Zhang, Yuhan Fu, Yihang Zhang +6 · 1 citation
    #cs.CL
  16. Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
    2026/07/27 by Tapan Parikh · 1 voice
    #cs.CL #cs.AI
  17. Kimi K3: Open Frontier Intelligence
    2026/07/27 by Kimi Team, Tongtong Bai, Yifan Bai +398 · 1 voice
    #cs.CL #cs.LG
  18. A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever
    2026/07/26 by Sietse Schelpe · 2 voices
    #cs.CL #cs.AI #cs.IR #cs.LG #cs.PF
  19. Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings
    2026/07/24 by Quentin Spencer · 2 voices
    #cs.CL
  20. The Boundaries of Automation: A Theory of Persistent Human Participation
    2026/07/23 by Fares Fourati, Hinrich Schütze, Eyke Hüllermeier +1 · 1 voice
    #cs.AI #cs.CL #cs.ET #cs.LG #cs.MA
  21. Error Certificates for KV-Cache Eviction via Randomized Design
    2026/07/23 by Peng Xie · 1 voice
    #cs.LG #cs.AI #cs.CL
  22. When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs
    2026/07/23 by Anna Mosolova, Djamé Seddah · 1 voice
    #cs.CL
  23. Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks
    2026/07/23 by Hiroki Tamba · 1 voice
    #cs.LG #cs.CL
  24. Agentic Evaluation of Copyright Law Compliance
    2026/07/23 by Zheng Hui, Doni Bloomfield, Noam Kolt · 1 citation
    #cs.CL #cs.CY
  25. Generative AI floods and dilutes the market for books
    2026/07/22 by Tuhin Chakrabarty, Xinyue Liu, Jane C. Ginsburg +1 · 3 voices
    #cs.CL #cs.AI #cs.CY
  26. Notes to Self: Can LLMs Benefit from Experiential Abstractions?
    2026/07/22 by Chang Liu, Xinyu Li, Artur Dubrawski · 1 voice
    #cs.CL
  27. Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking
    2026/07/22 by Kailin Jiang, Lei Liu, Jian Xi +7 · 1 citation
    #cs.CL
  28. surprisal is Not a Theory
    2026/07/22 by Andrés Buxó-Lugo, Aniello De Santo, Morgan Grobol +2 · 1 voice
    #cs.CL
  29. RAGAL: A Frugal, Fully Local Retrieval-Augmented Assistant for Technical Support at a Government Agency
    2026/07/21 by Dan Musetoiu · 1 voice
    #cs.IR #cs.CL
  30. Measuring Reward-Seeking via Contrastive Belief Updates
    2026/07/21 by Axel Højmark, Jérémy Scheurer, Evgenia Nitishinskaya +5 · 1 voice · 1 citation
    #cs.AI #cs.CL #cs.LG

more