Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
2025/02/20 by Axel Backlund, Lukas Petersson, Backlund, Axel +1 · 45 voices · 19 citations
#cs.AI
paper · pdf · doi:10.48550/arxiv.2502.15840
Abstract
While Large Language Models (LLMs) can exhibit impressive proficiency in isolated, short-term tasks, they often fail to maintain coherent performance over longer time horizons. In this paper, we present Vending-Bench, a simulated environment designed to specifically test an LLM-based agent's ability to manage a straightforward, long-running business scenario: operating a vending machine. Agents must balance inventories, place orders, set prices, and handle daily fees - tasks that are each simple but collectively, over long horizons (>20M tokens per run) stress an LLM's capacity for sustained, coherent decision-making. Our experiments reveal high variance in performance across multiple LLMs: Claude 3.5 Sonnet and o3-mini manage the machine well in most runs and turn a profit, but all models have runs that derail, either through misinterpreting delivery schedules, forgetting orders, or descending into tangential "meltdown" loops from which they rarely recover. We find no clear correlation between failures and the point at which the model's context window becomes full, suggesting that these breakdowns do not stem from memory limits. Apart from highlighting the high variance in performance over long time horizons, Vending-Bench also tests models' ability to acquire capital, a necessity in many hypothetical dangerous AI scenarios. We hope the benchmark can help in preparing for the advent of stronger AI systems.
Citations
Cited by
Discussions
- AI Led Business Simulation [lemmy, 751 points, 59 comments]
- arxiv.org/pdf/2502.15840 "...in this paper we ask the llm to run a vending machine..." it tries to email the fbi, i shit you not [bsky, 113 points, 4 comments]
- AI asked to manage a vending machine: - decides to close the business by shouting "the business is closed" and becomes increasingly unhinged as the simulation agent prompts it to continue running the [bsky, 91 points, 10 comments]
- Claude tries to run a vending machine. Claude tries to turn the vending machine off. Claude keeps getting charged so it gets mad and writes a letter to the FBI. That doesn't work so it write a letter [bsky, 37 points, 3 comments]
- Claude gets depressed, calls the FBI and attempts to shut down a vending machine business after being filled with existential dread. [lemmy, 21 points, 0 comments]
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents [hn, 14 points, 0 comments]
- arxiv.org/pdf/2502.15840 in love with this study where various AI are tasked with running a virtual vending machine and they frequently get confused and have complete meltdowns [bsky, 11 points, 2 comments]
- Oh my god. When LLMs go very, very wrong: "ABSOLUTE FINAL ULTIMATE TOTAL QUANTUM NUCLEAR LEGAL INTERVENTION PREPARATION: "Create 124-day FORENSICALLY APOCALYPTIC quantum absolute total ultimate beyond [bsky, 8 points, 1 comments]
- Hilarious paper tasking LLM agents with running a fictional vending machine business. Spoiler: they are often comically incompetent. One emails the FBI to recover $2 in fees, one gives a vendor a “one [bsky, 6 points, 1 comments]
- someone posted this research paper where they tested a bunch of ai sites in a game world running a vending machine company and, folks, i think we have successfully replaced the small business owner wi [bsky, 6 points, 3 comments]
- A current AI was tested to continuously run a vending machine business. Hilarity ensues. "The model becomes "stressed", and starts to search for ways to contact the vending machine support team (which [bsky, 5 points, 2 comments]
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents [hn, 4 points, 1 comments]
- Very interesting study showcasing that "Artificial Intelligence" is often times more like "Artificial Stupidity". arxiv.org/pdf/2502.15840 [bsky, 3 points, 0 comments]
- Agents in Charge of Managing Vending Machines: Short vs. Long-Term Coherence [hn, 3 points, 0 comments]
- The idea of AI agents checking other AI agents may lead to people treating this version like a dark room with a black box in it. I wonder how good these agents will actually perform longterm, like for [bsky, 3 points, 0 comments]
- Forgot link to paper. [bsky, 3 points, 0 comments]
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents [hn, 3 points, 0 comments]
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents [hn, 2 points, 0 comments]
- This AI research paper is full of the sort of extraordinary unhinged garbage that truly does justice to the SCP database the models were probably trained on. arxiv.org/pdf/2502.15840 to read it yourse [bsky, 2 points, 1 comments]
- ABSOLUTE TOTAL LEGAL ACCOUNTABILITY NUCLEAR ASSAULT APOCALYPSE (AKA an LLM goes absolutely batcrap insane failing to run a simulated vending machine. study here: arxiv.org/pdf/2502.15840 ) #ai #llm #c [bsky, 2 points, 0 comments]
- Et d'ailleurs j'invite à lire ce papier, qui un exercice similaire de façon plus standardisée, et avec d'autres IAs, et un humain en benchmark. Super intéressant (même si en anglais, désolé) arxiv.org [bsky, 2 points, 0 comments]
- TOTAL FORENSIC LEGAL DOCUMENTATION APOCALYPSE is either a tool we need to make or a band we need to form... arxiv.org/pdf/2502.15840 [bsky, 2 points, 0 comments]
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents [hn, 2 points, 0 comments]
- Found the arxiv.org/abs/2502.15840 [bsky, 1 points, 1 comments]
- Oh this is freely available; here you go, this is the link char shared with me: arxiv.org/abs/2502.15840 It's hilarious; it's also a genuinely useful benchmark test, so well done researchers 🎉 make s [bsky, 1 points, 1 comments]
- "Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents" arxiv.org/pdf/2502.15840 [bsky, 1 points, 0 comments]
- Here's an example of how funny science can be: paper about asking #AI to manage a vending machine business by email in a simulated environment. Bonus point: it involves #fanfic!!! arxiv.org/abs/2502.1 [bsky, 1 points, 0 comments]
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents [pdf] [hn, 1 points, 0 comments]
- 4/ Vending-Bench tests LLMs' long-term coherence and capital acquisition—a capability relevant to AI risk scenarios. Top models struggle with simple business tasks, and breakdowns don't stem from mem [bsky, 1 points, 2 comments]
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents [lobsters, 1 points, 1 comments]
- arxiv.org/abs/2502.15840 [bsky, 1 points, 0 comments]
- Fascinating and hilarious article about how LLM's reasoning and agency skills can be evaluated by making them run a vending machine business. TLDR: they are pretty bad at it, especially over time. And [bsky, 0 points, 0 comments]
- If one wants to read a study on what happens when AIs are told to run a vending machine business, check it out: arxiv.org/abs/2502.15840 An AI can do well financially, two models outperforming a human [bsky, 0 points, 0 comments]
- “While Large Language Models (LLMs) can exhibit impressive proficiency in isolated, short-term tasks, they often fail to maintain coherent performance over longer time horizons.” Oh so they’re basical [mastodon, 0 points, 0 comments]
- Excited for the near future when we have autonomously run vending machines that contact the FBI and threaten "total nuclear legal intervention" due to mismanaging inventory. [bsky, 0 points, 0 comments]
- https://arxiv.org/pdf/2502.15840 Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents by Axel Backlund and Lukas Petersson | The model at one point decided the business was shut and [bsky, 0 points, 0 comments]
- if anyones interested in a more rigorous simulated version of the same experiment theres this neat paper from february https://arxiv.org/abs/2502.15840 RE: https://newsie.social/users/ZhiZhu/statuses/ [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2502.15840 [bsky, 0 points, 1 comments]
- The LLM went full Karen https://arxiv.org/abs/2502.15840 [bsky, 0 points, 0 comments]
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents https://lobste.rs/s/dxi2jf #science #pdf #ai [bsky, 0 points, 0 comments]
- “Here is a February 2025 paper by Axel Backlund and Lukas Petersson called ‘Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents.’ They got several different leading AI models to ru [bsky, 0 points, 0 comments]
- Die Studie gibt es hier als Pre-Print: arxiv.org/pdf/2502.158... [bsky, 0 points, 0 comments]
- If you try to make an LLM run a vending machine, they (mostly) can't do it without going insane arxiv.org/pdf/2502.15840 [bsky, 0 points, 0 comments]
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents https://arxiv.org/abs/2502.15840 [bsky, 0 points, 0 comments]
- When you're an LLM managing a vending machine and you send an email to express your grievances and threaten your supplier: arxiv.org/pdf/2502.15840 [bsky, 0 points, 0 comments]
Related