LLMs Corrupt Your Documents When You Delegate
2026/04/17 by Philippe Laban, Tobias Schnabel, Jennifer Neville · 79 voices · 1 citation
#cs.CL #cs.HC
paper · pdf
Abstract
Large Language Models (LLMs) are poised to disrupt knowledge work, with the emergence of delegated work as a new interaction paradigm (e.g., vibe coding). Delegation requires trust - the expectation that the LLM will faithfully execute the task without introducing errors into documents. We introduce DELEGATE-52 to study the readiness of AI systems in delegated workflows. DELEGATE-52 simulates long delegated workflows that require in-depth document editing across 52 professional domains, such as coding, crystallography, and music notation. Our large-scale experiment with 19 LLMs reveals that current models degrade documents during delegation: even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of document content by the end of long workflows, with other models failing more severely. Additional experiments reveal that agentic tool use does not improve performance on DELEGATE-52, and that degradation severity is exacerbated by document size, length of interaction, or presence of distractor files. Our analysis shows that current LLMs are unreliable delegates: they introduce sparse but severe errors that silently corrupt documents, compounding over long interaction.
Citations
Cited by
Discussions
- LLMs corrupt your documents when you delegate [hn, 479 points, 201 comments]
- Microsoft researchers find that for non-programming tasks, LLM agents aren't ready for work. For long, truly complex work, they're only currently ready for Python programming. It's great that MS resea [bsky, 147 points, 6 comments]
- “LMs Corrupt Your Documents When You Delegate” https://arxiv.org/abs/2604.15597 > Our large-scale experiment with 19 LLMs reveals that current models degrade documents during delegation: even frontier [bsky, 44 points, 2 comments]
- LLMs Corrupt Your Documents When You Delegate [bsky, 34 points, 0 comments]
- Interesting new agent benchmark from MSR, DELEGATE-52, that "simulates long delegated workflows that require in-depth document editing across 52 professional domains". Main finding is that even fronti [bsky, 20 points, 1 comments]
- “Our large-scale experiment with 19 LLMs reveals that current models degrade documents during delegation: even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of d [bsky, 15 points, 0 comments]
- "We find that current LLMs are unreliable delegates: even frontier models corrupt an average of 25% of document content over long workflows, with sparse but severe errors that silently compound over t [bsky, 8 points, 0 comments]
- Microsoft Paper: LLMs Corrupt Your Documents When You Delegate (Arxiv.org) [hn, 7 points, 2 comments]
- 🤖 "Notre étude montre que les LLMs actuels sont des délégués peu fiables : ils introduisent des erreurs ponctuelles mais graves qui corrompent les documents en silence, et qui s'accumulent à force de [bsky, 5 points, 1 comments]
- LLMs Corrupt Your Documents When You Delegate [hn, 4 points, 0 comments]
- Interesting... "Our analysis shows that current LLMs are unreliable delegates: they introduce sparse but severe errors that silently corrupt documents, compounding over long interaction." Tonight's be [bsky, 4 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate [hn, 4 points, 2 comments]
- research arxiv.org/abs/2604.155... [bsky, 4 points, 0 comments]
- this table coming from Microsoft is interesting [bsky, 3 points, 0 comments]
- Unsurprising, I think. [bsky, 3 points, 0 comments]
- "... current LLMs are unreliable delegates: they introduce sparse but severe errors that silently corrupt documents, compounding over long interaction" Pre-print from Microsoft Research on LLM's abili [bsky, 3 points, 1 comments]
- LLMs Corrupt Your Documents When You Delegate [lemmy, 3 points, 0 comments]
- [2604.15597] LLMs Corrupt Your Documents When You Delegate arxiv.org/abs/2604.15597 - LLM이 인간을 대신해 복잡한 문서 편집 업무를 온전히 수행할 수 있는지 평가하기 위해, 52개 전문 분야를 포괄하는 시뮬레이션 벤치마크 'DELEGATE-52'를 개발 - 19개 모델을 테스트한 결과, [bsky, 2 points, 1 comments]
- LLMs Corrupt Your Documents When You Delegate Philippe Laban, Tobias Schnabel, Jennifer Neville arxiv.org/abs/2604.15597 github.com/microsoft/DE... [bsky, 2 points, 0 comments]
- We know that Ai is bobbins, but this paper calmly explains why: arxiv.org/pdf/2604.15597 No doubt that 25% will come down, but people are using this at scale NOW and don't realise the debt they are in [bsky, 2 points, 0 comments]
- These are the most dangerous errors, hard to spot but impactful long term: "current LLMs are unreliable delegates: they introduce sparse but severe errors that silently corrupt documents, compounding [bsky, 2 points, 2 comments]
- LLMs Corrupt Your Documents When You Delegate [hn, 2 points, 0 comments]
- Assuming by AI you mean LLM and associated agentic AI, no. Starter for 10: LLMs Corrupt Your Documents When You Delegate, Laban et al, 2026 arxiv.org/abs/2604.15597 Plus whose intellect was it trained [bsky, 1 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate https://arxiv.org/abs/2604.15597 (https://news.ycombinator.com/item?id=48073246) [bsky, 1 points, 0 comments]
- "Our large-scale experiment with 19 LLMs reveals that current models degrade documents during delegation: even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of d [bsky, 1 points, 1 comments]
- arxiv.org/abs/2604.15597 [bsky, 1 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate (Laban et al., 2026) arxiv.org/pdf/2604.15597 "[E]ven frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of document con [bsky, 1 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate. Large Language Models (LLMs) are poised to disrupt knowledge work, with the emergence of delegated work as a new interaction paradigm (e.g., vibe coding [bsky, 1 points, 0 comments]
- "Our large-scale experiment with 19 LLMs reveals that current models degrade documents during delegation: even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of d [bsky, 1 points, 1 comments]
- arxiv.org/abs/2604.15597 [bsky, 1 points, 0 comments]
- “[frontier models] losing on average 25% of document content over 20 delegated interactions, and an average degradation across all models of 50%.” Incredible. We were better off paying monks to hand-c [bsky, 1 points, 0 comments]
- Glad I don't use any AI agents, limit code generation to a few lines at most at any time, and verify sources for any question that goes beyond a Google search, because I care about the intentionality [bsky, 1 points, 0 comments]
- Yeah they're not even good for summarizing documents arxiv.org/pdf/2604.15597 [bsky, 1 points, 0 comments]
- [2/7] 📰 AI-generated LLMs can inadvertently corrupt documents when used to summarize or rewrite text, potentially leading to errors and inconsistencies. 🔗 https://arxiv.org/abs/2604.15597 #Tech #Dev [bsky, 1 points, 1 comments]
- The architectural takeaway: you can't treat LLMs as stateless editing oracles in stateful workflows. Explicit checkpointing, structural diffing, and domain-specific validation aren't nice-to-haves — t [bsky, 1 points, 0 comments]
- @davidgerard One for pivot: https://arxiv.org/pdf/2604.15597 Study by MS showing that even "frontier models" don't work to delegate tasks to because they fall apart. [bsky, 1 points, 0 comments]
- Need backup copies of every document you give to Agentic AI, I suppose: arxiv.org/abs/2604.15597 [bsky, 1 points, 0 comments]
- keep thinking about this paper which i saw because paco posted/shared it on linkedin a couple of days ago arxiv.org/abs/2604.15597 [bsky, 1 points, 1 comments]
- Good paper on document corruption with LLM’s. They are mean reversion machines. How do we make this more scalable, but mid? [bsky, 0 points, 0 comments]
- LLMs corrupt your documents when you delegate https:// arxiv.org/abs/2604.15597 # arxiv # llm # llms [mastodon, 0 points, 0 comments]
- LLMs corrupt your documents when you delegate https://arxiv.org/abs/2604.15597 https://news.ycombinator.com/item?id=48073246 [bsky, 0 points, 0 comments]
- [2604.15597] LLMs Corrupt Your Documents When You Delegate arxiv.org/abs/2604.15597 - LLM이 인간을 대신해 복잡한 문서 편집 업무를 온전히 수행할 수 있는지 평가하기 위해, 52개 전문 분야를 포괄하는 시뮬레이션 벤치마크 'DELEGATE-52'를 개발 - 19개 모델을 테스트한 결과, [bsky, 0 points, 1 comments]
- These are the most dangerous errors, hard to spot but impactful long term: "current LLMs are unreliable delegates: they introduce sparse but severe errors that silently corrupt documents, compounding [bsky, 0 points, 0 comments]
- Interesting abstract for "LLMs Corrupt Your Documents When You Delegate" I want to look into their testing and evaluation methods, but the fact that they are probing this area and mention multiple fac [bsky, 0 points, 0 comments]
- LLMs corrupt your documents when you delegate https://arxiv.org/abs/2604.15597 https://news.ycombinator.com/item?id=48073246 [bsky, 0 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate arxiv.org/pdf/2604.15597 [bsky, 0 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate #HackerNews https://arxiv.org/abs/2604.15597 [bsky, 0 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate [bsky, 0 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate "Our large-scale experiment with 19 LLMs reveals that current models degrade documents during delegation: even frontier models corrupt an average of 25% o [bsky, 0 points, 0 comments]
- Recent paper out of Cornell sheds more light on the state of task delegation to AI: "current LLMs are unreliable delegates: they introduce sparse but severe errors that silently corrupt documents, com [bsky, 0 points, 1 comments]
- [2604.15597] LLMs Corrupt Your Documents When You Delegate — Worth a read if you use LLMs to draft or edit shared docs: it shows how delegation can quietly introduce errors or unwanted changes. Good r [bsky, 0 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate https://arxiv.org/abs/2604.15597 (http://news.ycombinator.com/item?id=48073246) [bsky, 0 points, 0 comments]
- LLMs corrupt your documents when you delegate https://arxiv.org/abs/2604.15597 (http://news.ycombinator.com/item?id=48073246) [bsky, 0 points, 0 comments]
- So apparently we're just letting LLMs edit our work docs now, and half the time they're just corrupting the file instead. Guess I'm doing it myself. [bsky, 0 points, 0 comments]
- 2/2 la IA para tareas críticas de edición. Fuente: https://arxiv.org/abs/2604.15597 [bsky, 0 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate. Not a surprising outcome but good to have it on the record. [bsky, 0 points, 0 comments]
- LLMs corrupt documents in delegated work tasks. Badly. Even the best models. https://arxiv.org/abs/2604.15597 [bsky, 0 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate arxiv.org/pdf/2604.15597 [bsky, 0 points, 0 comments]
- Guess they'll have to go after microsoft for this unAmerican position paper, LLMs Corrupt Your Documents When You Delegate, arxiv.org/pdf/2604.15597 [bsky, 0 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate - Microsoft Research This would be some interesting reading [bsky, 0 points, 0 comments]
- "LLMs Corrupt Your Documents When You Delegate" arxiv.org/abs/2604.15597 [bsky, 0 points, 0 comments]
- It's amazing to me that Microsoft Research - a department of the documents and AI corp - has enough freedom to release this paper, which details how AI will mangle your documents over time (Interestin [bsky, 0 points, 1 comments]
- Ran across this pre-print exploring the tendency of LLM's to corrupt large datasets when delegated. It is not good for anyone for whom data quality is important. arxiv.org/pdf/2604.15597 [bsky, 0 points, 1 comments]
- Thanks. It reminded me of this paper from Microsoft research: arxiv.org/pdf/2604.15597 And this much struggle with a standardized dataset hints that in real life situations it will be worse, probably. [bsky, 0 points, 0 comments]
- Muy interesante sobre todo para las empresas: arxiv.org/pdf/2604.15597 básicamente no puedes confiar en que la IA con una instrucción sobre un documento no toque el resto del documento: LLMs Corrupt Y [bsky, 0 points, 1 comments]
- Paper in question: arxiv.org/abs/2604.15597 [bsky, 0 points, 1 comments]
- LLMs Corrupt Your Documents When You Delegate https://arxiv.org/abs/2604.15597 (https://news.ycombinator.com/item?id=48073246) [bsky, 0 points, 0 comments]
- it's also worth noting that microslop are trying to bury their own studies into how this shit hype-based tech VERY OFTEN CORRUPTS FILES WHEN GIVEN DELEGATED RESPONSIBILITY arxiv.org/abs/2604.15597 [bsky, 0 points, 0 comments]
- LLMs corrupt your documents when you delegate https://arxiv.org/abs/2604.15597 (https://news.ycombinator.com/item?id=48073246) [bsky, 0 points, 0 comments]
- Delegating tasks to LLMs can lead to unintended document corruption. Understanding the risks of relying on AI for writing and editing is essential. Stay informed to protect your work—let's discuss the [bsky, 0 points, 0 comments]
- LLMs corrupt your documents when you delegate https://arxiv.org/abs/2604.15597 comments #arxiv.org [bsky, 0 points, 0 comments]
- LLMs Corrupt Your Documents When You Delegate https://arxiv.org/abs/2604.15597 [bsky, 0 points, 0 comments]
- LLMs corrupt your documents when you delegate View Article | Join the HN Conversation Summary of HN discussion 🧵👇 [bsky, 0 points, 1 comments]
- LLMs corrupt your documents when you delegate https://arxiv.org/abs/2604.15597 [comments] [423 points] [bsky, 0 points, 0 comments]
- 📰 LLMs corrupt your documents when you delegate 🔗 https://arxiv.org/abs/2604.15597 💬 Discuss on HN [bsky, 0 points, 0 comments]
- AIs corrupt documents. I dont mean that by dismissing all of the actual content they write, but they inject errors into document they are working on, from unjustified deletions to hallucinations. Rese [bsky, 0 points, 1 comments]
- Zum Glück ist die Lösung dafür sehr einfach (do not delegate) [bsky, 0 points, 0 comments]
- arxiv.org/abs/2604.15597 Sounds about right [bsky, 0 points, 0 comments]
- Fascinating paper from MS research. Findings show that current LLMs introduce substantial errors when editing work documents, with frontier models losing on average 25% of document content over 20 del [bsky, 0 points, 1 comments]
Related