SWE-Lancer: Can Frontier LLMs Earn 1 Million from Real-World Freelance Software Engineering?
2025/02/17 by Samuel Miserendino, Miserendino, Samuel, Michele Wang +5 · 35 voices · 27 citations
#cs.LG #cs.SE
paper · pdf · doi:10.48550/arxiv.2502.12115
Abstract
We introduce SWE-Lancer, a benchmark of over 1,400 freelance software engineering tasks from Upwork, valued at $1 million USD total in real-world payouts. SWE-Lancer encompasses both independent engineering tasks--ranging from $50 bug fixes to $32,000 feature implementations--and managerial tasks, where models choose between technical implementation proposals. Independent tasks are graded with end-to-end tests triple-verified by experienced software engineers, while managerial decisions are assessed against the choices of the original hired engineering managers. We evaluate model performance and find that frontier models are still unable to solve the majority of tasks. To facilitate future research, we open-source a unified Docker image and a public evaluation split, SWE-Lancer Diamond (https://github.com/openai/SWELancer-Benchmark). By mapping model performance to monetary value, we hope SWE-Lancer enables greater research into the economic impact of AI model development.
Citations
Cited by
Discussions
- SWE-Lancer: a benchmark of freelance software engineering tasks from Upwork [hn, 111 points, 74 comments]
- OpenAI launched a new benchmark based on real-world software engineering tasks from Upwork Scores are awarded monetarily, by how much an AI could theoretically earn And Sonnet is currently the top m [bsky, 13 points, 3 comments]
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? [lobsters, 4 points, 3 comments]
- arxiv.org/abs/2502.12115 can't argue with the science on that one: LLM's are solving almost 60% of the manager tasks, but only 40% of SWE tasks :P [bsky, 2 points, 0 comments]
- arxiv.org/abs/2502.12115 [bsky, 2 points, 0 comments]
- OpenAI has a new benchmark where they run frontier LLMs against software development tasks on UpWork. So far, they fail to succeed on a majority of the possible $1 million payout from the tasks. But [bsky, 2 points, 1 comments]
- Time to ask for a pay rise I guess, OpenAI researches confirm that AI ain't going to replace developers any time soon: arxiv.org/pdf/2502.12115 [bsky, 2 points, 0 comments]
- paper not peer reviewed, probably never will be peer reviewed arxiv.org/pdf/2502.12115 [bsky, 1 points, 1 comments]
- SWE-Lancer is a new benchmark from OpenAI that evaluates large language models' ability to handle freelance software engineering tasks - Best model (Claude 3.5 Sonnet) only earned ~$400k out of possib [bsky, 1 points, 0 comments]
- I really like this idea from the SWE-Lancer paper where they give the agent a tool which simulates a user clicking through the app to verify that it works correctly. arxiv.org/abs/2502.12115 [bsky, 1 points, 0 comments]
- arxiv.org/pdf/2502.12115 [bsky, 1 points, 0 comments]
- https://bsky.app/profile/hackernews.com.web.brid.gy/post/3liigysdamrb2 [bsky, 0 points, 0 comments]
- https://bsky.app/profile/handle.invalid/post/3lii3q4oh422d [bsky, 0 points, 0 comments]
- SWE-Lancer: a benchmark of freelance software engineering tasks from Upwork https://arxiv.org/abs/2502.12115 [bsky, 0 points, 0 comments]
- SWE-Lancer: a benchmark of freelance software engineering tasks from Upwork https://arxiv.org/abs/2502.12115 [comments] [62 points] [bsky, 0 points, 0 comments]
- Can AI models do the tasks listed on Upwork? Evaluating tasks with a total remuneration of $1m, seems like the answer is no. Currently. arxiv.org/abs/2502.12115 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2502.12115 “Frontier models are still unable to solve the majority of tasks” . [bsky, 0 points, 0 comments]
- SWE-Lancer: a benchmark of freelance software engineering tasks from Upwork #HackerNews arxiv.org/abs/... [bsky, 0 points, 0 comments]
- Frontier AI models still struggle with real-world freelance software engineering. A new benchmark found that even the best model, Claude 3.5 Sonnet, earned less than half of $1M in available payouts—o [bsky, 0 points, 0 comments]
- SWE-Lancer: a benchmark of freelance software engineering tasks from Upwork [bsky, 0 points, 0 comments]
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? #dev #assistant #copilot #benchmark #llms #generativeai #upwork #freelance [bsky, 0 points, 0 comments]
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? - OpenAI introduce SWE-Lancer, a Benchmark of over 1,400 freelance Software Engineering Tasks from Upwork [bsky, 0 points, 0 comments]
- SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? https://lobste.rs/s/mswgiv #ai [bsky, 0 points, 0 comments]
- AI Native Industry Insights - Global - 20250220 - (2/4) SWE-Lancer: Benchmarking AI in Freelance Software Engineering Read more: https://arxiv.org/pdf/2502.12115 [bsky, 0 points, 0 comments]
- Si Claude Sonnet 3.5, la mejor IA en programación, fuera un freelance de Upwork, habría ganado 400K$ de 1M$ posibles, realizando tareas reales que se contratan a diario en la plataforma. Si quieres sa [bsky, 0 points, 0 comments]
- arxiv.org/abs/2502.12115 [bsky, 0 points, 1 comments]
- SWE-Lancer: a benchmark of freelance software engineering tasks from Upwork (arxiv.org) Main Link | Discussion [bsky, 0 points, 0 comments]
- Me dejo aquí el paper para echarle un vistazo mañana [bsky, 0 points, 1 comments]
- reading this paper: arxiv.org/abs/2502.12115 a benchmark with AI solving freelance software engineering tasks they do estimates on how much money using an LLM hypothetically saves over using a human ( [bsky, 0 points, 1 comments]
- OpenAI からSWE-Lancerっていう、割と実世界に近いと思われるフリーランスのプログラマーのタスクをこなすベンチマーク(ベンチマーク結果は何ドル稼いだか!)が公開されたんだけど… 結果: GPT-4o<o1<Claude 3.5 Sonnet 正直でよろしい :meow_lol: https://arxiv.org/abs/2502.12115 [bsky, 0 points, 1 comments]
- AI Can Write Code But Lacks Engineer's Instinct, OpenAI Study Finds Leading AI models can fix broken code, but they're nowhere near ready to replace human software engineers, according to extensive t [bsky, 0 points, 0 comments]
- AI is taking our jobs!.. or not. 🤷♂️ "SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?" arxiv.org/abs/2502.12115 Can't wait for an "update" of the pape [bsky, 0 points, 0 comments]
- Test harness for LLM agents for carrying out real-world software engineering tasks. The upshot is that autonomous software development remains “challenging” for even frontier models. https://arxiv.or [bsky, 0 points, 0 comments]
- Novel benchmark testing if top LLMs can handle $1M worth of real Upwork coding tasks. Spoiler: They can't. Yet. Worth tracking because when models start passing this, software engineering jobs get i [bsky, 0 points, 0 comments]
- AI belonging to Anthropic, who's CEO penned the optimistic 'Machines of Loving Grace', just automated away 40% of software engineering work on a leading freelancer platform. [lemmy, -2 points, 2 comments]
Related