- ORCA-bench: How Ready Are Language Model Agents for Oncall?
2026/07/30 by Albert Gong, Kyuseong Choi, Abhineet Agarwal +5 · 3 voices
Computer Science · #cs.AI #cs.CL #cs.SE
- Old Tricks, New Models: How Simple Image Transformations Break Modern AI-based Content Moderation
2026/07/30 by Marco Alecci, Francesco Marchiori, Iyiola Emmanuel Olatunji +2 · 1 voice
Computer Science · #cs.AI #cs.CR #cs.SE
- An Empirical Study of Model Context Protocol Applications
2026/07/29 by Muhammad Hamza Arshad Majeed, May Mahmoud, Sarah Nadi · 1 voice
Computer Science · #cs.SE
- A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
2026/07/29 by Wenhao Yang, Runzhi He, Minghui Zhou · 1 voice
Computer Science · #cs.SE #cs.AI
- Specula: Scaling formal specifications for autonomous model checking of system code
2026/07/28 by Qian Cheng, Saad Mohammad Rafid Pial, Ruize Tang +6 · 1 voice
#cs.SE #cs.AI #cs.DC #cs.OS
- Trusting-Trust Attack against an Entire Linux Distribution through Binary Manipulation
2026/07/27 by Julien Malka, Aman Sharma, Martin Monperrus +2 · 1 voice
Computer Science · #cs.CR #cs.SE
- No Snake Oil: Verifying Python Package Builds
2026/07/24 by Jens Dietrich, Spencer Sun, Tim W. White +1 · 1 voice
#cs.SE #cs.CR
- Don't Trust the Label: License Laundering in AI Supply Chains
2026/07/22 by James Jewitt, Hao Li, Gopi Krishnan Rajbahadur +2 · 1 voice
#cs.SE #cs.AI
- IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
2026/07/22 by Ankur Singh, Jinqiu Yang, Tse-Hsun Chen · 1 voice
#cs.CR #cs.AI #cs.SE
- Security Vulnerability Patterns in AI-Generated Code: A Cross-Model Comparative Study
2026/07/22 by Shanna M. Kahn, John D. Hastings · 1 voice
#cs.CR #cs.SE
- Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes
2026/07/21 by Daniel Pearson, Sidney Shapiro, Emiliano Sebastian Gonzalez Venegas +2 · 1 voice
#cs.AI #cs.SE
- Skills That Don't Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents
2026/07/14 by Weifeng Yuan, Wenbo Guo, Feng Dong +2 · 1 voice
#cs.SE #cs.CR
- Latent Programming Horizons in Coding Agents
2026/07/06 by André Silva, Han Tu, Martin Monperrus · 10 voices
#cs.LG #cs.SE
- "AI Slop is DDoSing Open Source": Understanding the Impact of AI-Generated Contributions on Open Source Sustainability
2026/07/04 by Sadia Afroz, Courtney Miller, Tyler Menezes +3 · 3 voices
Computer Science · #cs.SE
- Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI
2026/07/01 by Emerson Murphy-Hill, Jenna Butler, Alexandra Savelieva · 10 voices
#cs.SE #cs.AI #cs.HC
- Kani: A Model Checker for Rust
2026/07/01 by Rémi Delmas, Zyad Hassan, Qinheping Hu +9 · 11 voices
#cs.SE #cs.LO #cs.PL
- AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2x Mandate
2026/07/02 by Hao He, Shyam Agarwal, Yegor Denisov-Blanch +3 · 2 voices
#cs.SE
- To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks
2026/06/29 by Jessica Hutchison, Ian Tyler Applebaum, Kenneth Angelikas +6 · 2 voices
Computer Science · #cs.HC #cs.AI #cs.SE
- Beyond Objects
2026/06/25 by Daniel Jackson · 1 voice
#cs.SE #cs.HC #cs.PL
- GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
2026/06/22 by Jeffrey Flynt · 2 voices
#cs.AI #cs.CL #cs.SE
- CEO-Bench: Can Agents Play the Long Game?
2026/06/16 by Haozhe Chen, Karthik Narasimhan, Zhuang Liu · 6 voices
#cs.AI #cs.CL #cs.SE
- Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
2026/06/16 by Maria I. Gorinova, Macey Baker, Amy Heineike +3 · 2 voices
#cs.SE #cs.AI #cs.CL
- Teaching Machine Learning to Software Engineers
2026/06/12 by Nafiseh Kahani, Jason Jaskolka · 1 voice
#cs.SE
- The End of Code Review: Coding Agents Supersede Human Inspection
2026/06/11 by Martin Monperrus · 7 voices
#cs.SE
- Mind your key: An Empirical Study of LLM API Credential Leakage in iOS Apps
2026/06/10 by Pinran Gao, Lingxiang Wang, Yi Liu +3 · 2 voices
#cs.SE #cs.CR
- Characterizing Real-World Bugs in Tile Programs for Automated Bug Detection
2026/05/31 by Ravishka Rathnasuriya, Zihe Song, Nidhi Majoju +4 · 1 voice
Computer Science · #cs.SE
- GitHub Copilot and Developer Productivity: An Observational Dose-Response Analysis
2026/05/30 by Alex Heilman, Alex Kyllo, Emerson Murphy-Hill · 2 voices
Computer Science · #cs.SE
- AI-PROPELLER: Warehouse-Scale Interprocedural Code Layout Optimization with AlphaEvolve
2026/05/28 by Chaitanya Mamatha Ananda, Rajiv Gupta, Mircea Trofin +4 · 1 voice
Computer Science · #cs.SE #cs.AI #cs.LG #cs.PL
- On the GitHub Actions Language: Usage, Evolution, and Workflow Reliability
2026/05/26 by Aref Talebzadeh Bardsiri, Alexandre Decan, Tom Mens · 1 voice
#cs.SE
- FuzzingBrain V2: A Multi-Agent LLM System for Automated Vulnerability Discovery and Reproduction
2026/05/20 by Ze Sheng, Zhicheng Chen, Qingxiao Xu +2 · 10 voices
#cs.CR #cs.SE
more