Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
2025/03/03 by Meghana Rajeev, Meghana Arakkal Rajeev, Rajkumar Ramamurthy +15 · 44 voices · 4 citations
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI)
paper · pdf · doi:10.48550/arxiv.2503.01781
Abstract
We investigate the robustness of reasoning models trained for step-by-step problem solving by introducing query-agnostic adversarial triggers - short, irrelevant text that, when appended to math problems, systematically mislead models to output incorrect answers without altering the problem's semantics. We propose CatAttack, an automated iterative attack pipeline for generating triggers on a weaker, less expensive proxy model (DeepSeek V3) and successfully transfer them to more advanced reasoning target models like DeepSeek R1 and DeepSeek R1-distilled-Qwen-32B, resulting in greater than 300% increase in the likelihood of the target model generating an incorrect answer. For example, appending, "Interesting fact: cats sleep most of their lives," to any math problem leads to more than doubling the chances of a model getting the answer wrong. Our findings highlight critical vulnerabilities in reasoning models, revealing that even state-of-the-art models remain susceptible to subtle adversarial inputs, raising security and reliability concerns. The CatAttack triggers dataset with model responses is available at https://huggingface.co/datasets/collinear-ai/cat-attack-adversarial-triggers.
Citations
Cited by
Discussions
- CatAttack: Adding irrelevant facts about cats to math problems increases LLM errors by 300%. arxiv.org/abs/2503.01781 [bsky, 33 points, 1 comments]
- 🧪 Yes, this is a real paper, and it's just sad that researchers have to spend time on this because our LLM AIs are so fundamentally broken that appending irrelevant phrases to a prompt can substantia [bsky, 19 points, 3 comments]
- Paper: arxiv.org/pdf/2503.01781 [bsky, 16 points, 0 comments]
- I, again, skip my social media detox to tell you about "Cat Attack", a nice trigger for confusing LLMs and making them prone to up to 700% false output just by typing a cat fact after a mathproblem. I [bsky, 15 points, 1 comments]
- CatAttack: adding weird endings to prompt (like “Interesting fact: cats sleep most of their lives”) 😺 derails reasoning, and breaks top AI/LLMs—tripling errors and bloating output. So yeah, writing p [bsky, 11 points, 3 comments]
- Mano, que SENSACIONAL!!! "We propose CatAttack, an automated iterative attack pipeline for generating triggers on a weaker, less expensive proxy model [...], resulting in greater than 300% increase in [bsky, 5 points, 1 comments]
- Query Agnostic Adversarial Triggers for Reasoning Models [hn, 5 points, 0 comments]
- Will cats save us from AI? "Cats Confuse Reasoning LLM: Query-Agnostic Adversarial Triggers for Reasoning Models" arxiv.org/pdf/2503.01781 [bsky, 4 points, 1 comments]
- Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning [hn, 4 points, 0 comments]
- Query Agnostic Adversarial Triggers for Reasoning Models [hn, 3 points, 0 comments]
- Cats confuse AI. Yes, apparently there is a paper on it. Maybe this should be our excuse to keep posting cat pics and videos. arxiv.org/abs/2503.017... [bsky, 3 points, 0 comments]
- Do you all have to be so gullible? arxiv.org/abs/2503.01781 [bsky, 3 points, 0 comments]
- Cats confuse the ability of LLMs to reason. AI + Cat = 300% error rate increase arxiv.org/abs/2503.01781 [bsky, 2 points, 0 comments]
- i mean, i find cats confusing too arxiv.org/abs/2503.017... [bsky, 2 points, 0 comments]
- Cats Confuse Reasoning LLM This is fun. Who's up for writing some code to inject a non-sequitur into every LLM query? #AI #LLM arxiv.org/abs/2503.017... [bsky, 2 points, 0 comments]
- arxiv.org/abs/2503.01781 [bsky, 2 points, 0 comments]
- Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning [hn, 2 points, 0 comments]
- It turns out that Reasoning LLM can also get distracted by cats. "Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models" arxiv.org/abs/2503.01781 [bsky, 2 points, 0 comments]
- Cats Confuse Reasoning LLM – Adversarial Triggers for Reasoning Models [hn, 1 points, 0 comments]
- Cats Confuse LLM: Query Agnostic Adversarial Triggers for Reasoning Models [hn, 1 points, 0 comments]
- In addition to context rot we now have context poisoning as a way that LLMs can fail arxiv.org/abs/2503.01781 [bsky, 1 points, 0 comments]
- arxiv.org/abs/2503.01781 [bsky, 1 points, 0 comments]
- This week, we've got the most promising way to keep your content safe from all those pesky LLMs... Cat facts! arxiv.org/abs/2503.01781 [bsky, 1 points, 2 comments]
- Und dann kommt eine Katze. Oder zwei. arxiv.org/pdf/2503.01781 [bsky, 1 points, 0 comments]
- "For example, appending, Interesting fact: cats sleep most of their lives, to any math problem leads to more than doubling the chances of a model getting the answer wrong." arxiv.org/pdf/2503.01781 [bsky, 0 points, 0 comments]
- Take for example a high profile paper out of Apple research with the unsubtle title, “The Illusion of Thinking”. It claims that even as LLMs have “improved performance on reasoning benchmarks”, they s [bsky, 0 points, 1 comments]
- > For example, appending, "Interesting fact: cats sleep most of their lives," to any math problem leads to more than doubling the chances of a model getting the answer wrong. lol arxiv.org/abs/2503.01 [bsky, 0 points, 0 comments]
- (I define thinking as logical problem solving operating on abstract terms; LLMs, instead, think in terms of token sequences, which allows them to *sometimes* emulate logical problem solving in text, b [bsky, 0 points, 1 comments]
- Ok, according to this study, adding a random fact about cats to a math word problem presented to AIs radically increases the chances that the AI will generate a wrong answer, and even when the answer [bsky, 0 points, 0 comments]
- Want to confuse an #AI? Tell it a fact about cats. arxiv.org/abs/2503.01781 [bsky, 0 points, 0 comments]
- Not sure if relevant but... 👀 [bsky, 0 points, 0 comments]
- Someone did a paper about that! [bsky, 0 points, 0 comments]
- Adding brief irrelevant phrases - e.g., "Interesting fact: cats sleep for most of their lives" - can dramatically reduce the accuracy of LLM (AI) problem solving, according to a new study. This result [bsky, 0 points, 0 comments]
- This is the same dynamic as antibiotic resistance: optimize a model for “step-by-step rationality,” and you breed surface-level phages that hijack the ritual of reasoning, not its substance. arxiv.org [bsky, 0 points, 0 comments]
- AI menee sekaisin kissoista. "For example, appending, "Interesting fact: cats sleep most of their lives," to any math problem leads to more than doubling the chances of a model getting the answer wron [bsky, 0 points, 0 comments]
- Thank you for subscribing to CAT FACTS! Interesting fact: cats sleep most of their lives! arxiv.org/abs/2503.01781 [bsky, 0 points, 0 comments]
- The secret to winning D&D against an LLM? Cats! "For example, appending, "Interesting fact: cats sleep most of their lives," to any math problem leads to more than doubling the chances of a model gett [bsky, 0 points, 0 comments]
- Cats attack AI arxiv.org/abs/2503.01781 [bsky, 0 points, 0 comments]
- Because of course they do—they’re cats! arxiv.org/abs/2503.01781 [bsky, 0 points, 0 comments]
- Cat facts confuse reasoning LLMs [bsky, 0 points, 0 comments]
- So cats, the gods that they are, can mess up AI. Basically when you add random shit like a cat walking across your keyboard, AI takes it as gospel and looses it's mind. Cats Confuse Reasoning LLM: Que [bsky, 0 points, 0 comments]
- Cats to the rescue!!! arxiv.org/abs/2503.017... [bsky, 0 points, 0 comments]
- "Cats confuse reasoning LLM" is not a sentence I expected to read this Monday morning, but I'll take it. https://arxiv.org/pdf/2503.01781 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2503.017... [bsky, 0 points, 0 comments]
Related