vix.ing · top · new · best · stats · spec
  1. ORCA-bench: How Ready Are Language Model Agents for Oncall?
    2026/07/30 by Albert Gong, Kyuseong Choi, Abhineet Agarwal +5 · 3 voices
    Computer Science · #cs.AI #cs.CL #cs.SE
  2. MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems
    2026/07/30 by Mao-xun Huang, Jerry Wang, Yi-Cheng Lai +3 · 1 voice
    Computer Science · #cs.AI
  3. Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B
    2026/07/30 by Iliya Mirzaei · 1 voice
    Computer Science · #cs.AI #cs.CL #cs.LG
  4. AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
    2026/07/30 by Xiangning Lin, Shenzhe Zhu, Shu Yang +23 · 1 voice
    Computer Science · #cs.AI #cs.CL #cs.CY #cs.HC
  5. AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
    2026/07/30 by Bing Yan, Gregory Wolfe, Stefano Martiniani +1 · 1 voice
    Computer Science · #cs.AI #cs.CL #cs.IR #cs.LG
  6. Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation
    2026/07/30 by Marco Alecci, Francesco Marchiori, Iyiola Emmanuel Olatunji +2 · 1 voice
    Computer Science · #cs.AI #cs.CR #cs.SE
  7. Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
    2026/07/30 by Hanzhang Zhou, Panrong Tong, Xu Zhang +13 · 1 voice
    Computer Science · #cs.AI #cs.CV
  8. OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
    2026/07/30 by Qiushi Sun, Kanzhi Cheng, Yian Wang +20 · 1 voice
    Computer Science · #cs.AI #cs.CL #cs.CV
  9. Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
    2026/07/29 by Alexi Gladstone, Heng Ji, Yilun Du · 1 voice
    Computer Science · #cs.LG #cs.AI #cs.CL #cs.CV
  10. HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
    2026/07/28 by Liudas Panavas, Sebastian Minus, Bradley Monton +4 · 3 voices
    Computer Science · #cs.AI #cs.CL
  11. Can AI agents conduct open-ended AI research? Early evidence from two case studies
    2026/07/29 by Peter Kirgis, Sayash Kapoor, Andrew Schwartz +21 · 2 voices
    Computer Science · #cs.AI #cs.CY #cs.LG
  12. Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction
    2026/07/29 by Xin Xu, Chengrui Wu, Jiayu Lu +3 · 2 voices
    Computer Science · #cs.AI #cs.CR #cs.GT
  13. A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
    2026/07/29 by Wenhao Yang, Runzhi He, Minghui Zhou · 1 voice
    Computer Science · #cs.SE #cs.AI
  14. Hearsay: Vision-Language Medical Diagnoses Without an Image
    2026/07/29 by Siddharth Vohra · 1 voice
    Computer Science · #cs.CV #cs.AI #cs.CL #cs.CY
  15. Human diversity fuels collective creativity that large language models cannot simulate or sustain
    2026/07/29 by Mengchen Dong, Hiromu Yakura · 1 voice
    Computer Science · #cs.HC #cs.AI #cs.CY
  16. Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
    2026/07/28 by Jiabao Ji, Yujian Liu, Li An +4 · 1 voice
    #cs.AI
  17. Visual prompt engineering for video models
    2026/07/28 by Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer +7 · 1 voice
    #cs.CV #cs.AI
  18. Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
    2026/07/28 by Stefan Krsteski, Charlotte Meyer, Guillaume Allegre +2 · 1 voice
    #cs.AI #cs.DB
  19. Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors
    2026/07/28 by Jianfei Ma, Zhaoxin Feng, Emmanuele Chersoni +1 · 1 voice
    Computer Science · #cs.AI #cs.CL
  20. Specula: Scaling formal specifications for autonomous model checking of system code
    2026/07/28 by Qian Cheng, Saad Mohammad Rafid Pial, Ruize Tang +6 · 1 voice
    #cs.SE #cs.AI #cs.DC #cs.OS
  21. At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference
    2026/07/28 by Bowen Wang, Chi Zhang, Diyou Shen +3 · 1 voice
    #cs.AR #cs.AI
  22. A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series
    2026/07/28 by Frank Nie, Ethan B Liu, Yuan Zhu +2 · 1 citation
    #cs.AI #cs.CL
  23. From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
    2026/07/27 by Junlin Liu, Jiangwang Chen, Zixin Song +7 · 2 voices · 1 citation
    #cs.AI
  24. Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
    2026/07/27 by Tapan Parikh · 1 voice
    #cs.CL #cs.AI
  25. Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents
    2026/07/27 by Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov +2 · 1 voice
    #cs.CR #cs.AI
  26. The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing
    2026/07/27 by Stefan Scholze, Johannes Partzsch, Sebastian Höppner +27 · 1 voice
    #cs.ET #cs.AI #cs.AR #cs.DC
  27. A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever
    2026/07/26 by Sietse Schelpe · 2 voices
    #cs.CL #cs.AI #cs.IR #cs.LG #cs.PF
  28. An Exact Counterexample to Carlson's Associated-Prime Depth Conjecture from a Group of Order 128
    2026/07/26 by Xinan Dai, Wenhao Deng, Yingdong Shi +2 · 1 citation
    #math.GR #cs.AI
  29. Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI
    2026/07/24 by Jiaqi Shao, Hanck Chen, Wei Zhang +2 · 1 voice
    #cs.AI
  30. Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
    2026/07/23 by Gaurav Dadhich · 2 voices
    #cs.AI #cs.IR

more