Large Language Models are Zero-Shot Reasoners
2022/05/24 by Takeshi Kojima, Kojima, Takeshi, Shixiang Gu +8 · 10 voices · 713 citations
Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2205.11916
openalex publication_date 2022/05/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Pretrained large language models (LLMs) are widely used in many sub-fields of natural language processing (NLP) and generally known as excellent few-shot learners with task-specific exemplars. Notably, chain of thought (CoT) prompting, a recent technique for eliciting complex multi-step reasoning through step-by-step answer examples, achieved the state-of-the-art performances in arithmetics and symbolic reasoning, difficult system-2 tasks that do not follow the standard scaling laws for LLMs. While these successes are often attributed to LLMs' ability for few-shot learning, we show that LLMs are decent zero-shot reasoners by simply adding "Let's think step by step" before each answer. Experimental results demonstrate that our Zero-shot-CoT, using the same single prompt template, significantly outperforms zero-shot LLM performances on diverse benchmark reasoning tasks including arithmetics (MultiArith, GSM8K, AQUA-RAT, SVAMP), symbolic reasoning (Last Letter, Coin Flip), and other logical reasoning tasks (Date Understanding, Tracking Shuffled Objects), without any hand-crafted few-shot examples, e.g. increasing the accuracy on MultiArith from 17.7% to 78.7% and GSM8K from 10.4% to 40.7% with large InstructGPT model (text-davinci-002), as well as similar magnitudes of improvements with another off-the-shelf large model, 540B parameter PaLM. The versatility of this single prompt across very diverse reasoning tasks hints at untapped and understudied fundamental zero-shot capabilities of LLMs, suggesting high-level, multi-task broad cognitive capabilities may be extracted by simple prompting. We hope our work not only serves as the minimal strongest zero-shot baseline for the challenging reasoning benchmarks, but also highlights the importance of carefully exploring and analyzing the enormous zero-shot knowledge hidden inside LLMs before crafting finetuning datasets or few-shot exemplars.
Cited by
- CORE: A Unified Cascaded Ordinal Relevance Estimation Framework for E-commerce Search
- CME-CAD: Heterogeneous Collaborative Multi-Expert Reinforcement Learning for CAD Code Generation
- Robust LLM-based Column Type Annotation via Prompt Augmentation with LoRA Tuning
- Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- Mitigating Social Desirability Bias in Random Silicon Sampling
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- Chain-of-thought Reviewing and Correction for Time Series Question Answering
- Temporal Visual Semantics-Induced Human Motion Understanding with Large Language Models
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Semantic-Enhanced Automatic Refinement of Architecture Recovery Results Using LLMs
- Separating Clicks from Baits: Using Large Language Models to Detect Misleading YouTube Thumbnails
- Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation
- Verbalized Particle Posterior: Bayesian Inference over Natural Language Hypotheses
- Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding
- SymStep: Symbolic Step Verification for Logical Reasoning
- Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
- Design Theater: A Benchmark for Generative UI
- Toward Human-Centered Multi-Agent Systems: Integrating Cognition, Culture, Values, and Cooperation in AI Agents
- Some hypotheses on how chatbots work in problem-solving-driven conversations. Large Language Models as confirmation of the Innovation Illusion
- Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn
- Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering
- Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning
- What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation
- Explainable Statute Prediction via Attention-based Model and LLM Prompting
- Exploring the Heterogeneity of Tabular Data: A Diversity-aware Data Generator via LLMs
- Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models
- Method Decoration (DeMe): A Framework for LLM-Driven Adaptive Method Generation in Dynamic IoT Environments
- Do Latent Tokens Think? A Causal and Adversarial Analysis of Chain-of-Continuous-Thought
- Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
- Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models
- BRIDGE: Budget-aware Reasoning via Intermediate Distillation with Guided Examples
- Learning to Reason in LLMs by Expectation Maximization
- LoFT-LLM: Low-Frequency Time-Series Forecasting with Large Language Models
- The Mental World of Large Language Models in Recommendation: A Benchmark on Association, Personalization, and Knowledgeability
- Observer, Not Player: Simulating Theory of Mind in LLMs through Game Observation
- Can abstract concepts from LLM improve SLM performance?
- Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing
- Auto-Prompting with Retrieval Guidance for Frame Detection in Logistics
- CienaLLM: Generative Climate-Impact Extraction from News Articles with Autoregressive LLMs
- Beyond the Prompt: An Empirical Study of Cursor Rules
- Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs
- MDToC: Metacognitive Dynamic Tree of Concepts for Boosting Mathematical Problem-Solving of Large Language Models
- LLMs on Drugs: Language Models Are Few-Shot Consumers
- External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
- Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
- Towards Deeper Emotional Reflection: Crafting Affective Image Filters with Generative Priors
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
- A systematic assessment of Large Language Models for constructing two-level fractional factorial designs
- BRAID: Bounded Reasoning for Autonomous Inference and Decisions
- Explaining the Reasoning of Large Language Models Using Attribution Graphs
- Dual-Density Inference for Efficient Language Model Reasoning
- Evaluating LLMs for Zeolite Synthesis Event Extraction (ZSEE): A Systematic Analysis of Prompting Strategies
- Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
- Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
- ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
- C-ing Clearly: Enhanced Binary Code Explanations using C code
- Georeferencing complex relative locality descriptions with large language models
- LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation
- Intention Chain-of-Thought Prompting with Dynamic Routing for Code Generation
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Do Reviews Matter for Recommendations in the Era of Large Language Models?
- State over Tokens: Characterizing the Role of Reasoning Tokens
- How Prompts Move Language Model Behavior: Frames, Salience, and Construal as Semantic Control
- Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
- Taint-Based Code Slicing for LLMs-based Malicious NPM Package Detection
- Instruction-Tuning Open-Weight Language Models for BPMN Model Generation
- Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs
- Designing AI-Resilient Assessments Using Interconnected Problems: A Theoretically Grounded and Empirically Validated Framework
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- Exploring LLMs for Scientific Information Extraction Using The SciEx Framework
- Understanding Chain-of-Thought Effectiveness in Code Generation: An Empirical and Information-Theoretic Analysis
- Rethinking Chain-of-Thought Reasoning for Videos
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- A Multi-Robot Platform for Robotic Triage Combining Onboard Sensing and Foundation Models
- To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
- Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation
- Unified Video Editing with Temporal Reasoner
- Attention as Binding: A Vector-Symbolic Perspective on Transformer Reasoning
- Block Sparse Flash Attention
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code
- PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
- The Road of Adaptive AI for Precision in Cybersecurity
- Concept-based Explainable Data Mining with VLM for 3D Detection
- Empirical Prompt Engineering for Construct Identification with Large Language Models
- AR-Med: Automated Relevance Enhancement in Medical Search via LLM-Driven Information Augmentation
- Understanding LLM Reasoning for Abstractive Summarization
- See, Think, Learn: A Self-Taught Multimodal Reasoner
- DETAIL Matters: Measuring the Impact of Prompt Specificity on Reasoning in Large Language Models
- Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
- Improving Phishing Resilience with AI-Generated Training: Evidence on Prompting, Personalization, and Duration
- Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
- Zero-Overhead Introspection for Adaptive Test-Time Compute
- ReJump: A Tree-Jump Representation for Analyzing and Improving LLM Reasoning
- T2T-LA: A Topology-to-Topology LLM Agent for Graph Learning with Neither Feature Access nor Task Knowledge
- Demystifying Errors in LLM Reasoning Traces: An Empirical Study of Code Execution Simulation
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- Toward Automatic Safe Driving Instruction: A Large-Scale Vision Language Model Approach
- Listwise Preference Optimization with Element-wise Confusions for Aspect Sentiment Quad Prediction
- An Empirical Study on the Security Vulnerabilities of GPTs
- Multi-chain Graph Refinement and Selection for Reliable Reasoning in Large Language Models
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- Analyzing Image Beyond Visual Aspect: Image Emotion Classification via Multiple-Affective Captioning
- AI-Generated Compromises for Coalition Formation: Modeling, Simulation, and a Textual Case Study
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- RecToM: A Benchmark for Evaluating Machine Theory of Mind in LLM-based Conversational Recommender Systems
- Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
- On the Limits of Innate Planning in Large Language Models
- Chatty-KG: A Multi-Agent AI System for On-Demand Conversational Question Answering over Knowledge Graphs
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- LLM-Driven Transient Stability Assessment: From Automated Simulation to Neural Architecture Design
- More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering
- A Machine Learning Approach for Detection of Mental Health Conditions and Cyberbullying from Social Media
- The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models
- Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
- Cross Domain Evaluation of Multimodal Chain-of-Thought Reasoning of different datasets into the Amazon CoT Framework
- Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment
- GraphMind: Theorem Selection and Conclusion Generation Framework with Dynamic GNN for LLM Reasoning
- Skeletons Matter: Dynamic Data Augmentation for Text-to-Query
- HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- Eliciting Chain-of-Thought in Base LLMs via Gradient-Based Representation Optimization
- HuggingR4: A Progressive Reasoning Framework for Discovering Optimal Model Companions
- Prompt Optimization as a State-Space Search Problem
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- Efficient Robot Design with Multi-Objective Black-Box Optimization and Large Language Models
- FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle
- Hallucinate Less by Thinking More: Aspect-Based Causal Abstention for Large Language Models
- ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
- HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor Generation
- ReVul-CoT: Towards Effective Software Vulnerability Assessment with Retrieval-Augmented Generation and Chain-of-Thought Prompting
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- ELPO: Ensemble Learning Based Prompt Optimization for Large Language Models
- GazeInterpreter: Parsing Eye Gaze to Generate Eye-Body-Coordinated Narrations
- PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- As If We've Met Before: LLMs Exhibit Certainty in Recognizing Seen Files
- SafeRBench: A Comprehensive Benchmark for Safety Assessment in Large Reasoning Models
- GPS: General Per-Sample Prompter
- M-CALLM: Multi-level Context Aware LLM Framework for Group Interaction Prediction
- Dynamic Template Selection for Output Token Generation Optimization: MLP-Based and Transformer Approaches
- Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- Genomic Next-Token Predictors are In-Context Learners
- Advanced Tool for Traffic Crash Analysis: An AI-Driven Multi-Agent Approach to Pre-Crash Reconstruction
- Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations
- Look as You Think: Unifying Reasoning and Visual Evidence Attribution for Verifiable Document RAG via Reinforcement Learning
- Scaling Open-Weight Large Language Models for Hydropower Regulatory Information Extraction: A Systematic Analysis
- On the Notion that Language Models Reason
- Free3D: 3D Human Motion Emerges from Single-View 2D Supervision
- On the Measure of a Model: From Intelligence to Generality
- Analogical Structure, Minimal Contextual Cues and Contrastive Distractors: Input Design for Sample-Efficient Linguistic Rule Induction
- Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
- Efficient Thought Space Exploration Through Strategic Intervention
- ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
- PustakAI: Curriculum-Aligned and Interactive Textbooks Using Large Language Models
- DemoTuner: Efficient DBMS Knobs Tuning via LLM-Assisted Demonstration Reinforcement Learning
- Mastering Olympiad-Level Physics with Artificial Intelligence
- Towards Effective and Efficient Non-autoregressive decoders for Conformer and LLM-based ASR using Block-based Attention Mask
- Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
- HalluClean: A Unified Framework to Combat Hallucinations in LLMs
- Temporal Predictors of Outcome in Reasoning Language Models
- Representation Learning Enables Scalable Multitask Deep Reinforcement Learning
- General Intelligence-based Fragmentation (GIF): A framework for peak-labeled spectra simulation
- Still Not There: Can LLMs Outperform Smaller Task-Specific Seq2Seq Models on the Poetry-to-Prose Conversion Task?
- Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- Beyond Correctness: Evaluating and Improving LLM Feedback in Statistical Education
- On the Creativity of AI Agents
- SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
- S-DAG: A Subject-Based Directed Acyclic Graph for Multi-Agent Heterogeneous Reasoning
- FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language Models
- Rank-1 LoRAs Encode Interpretable Reasoning Signals
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models
- Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
- An Empirical Study of Reasoning Steps in Thinking Code LLMs
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
- Effectiveness of Chain-of-Thought in Distilling Reasoning Capability from Large Language Models
- What About Our Bug? A Study on the Responsiveness of NPM Package Maintainers
- Logit-Entropy Adaptive Stopping Heuristic for Efficient Chain-of-Thought Reasoning
- Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
- Large language models replicate and predict human cooperation across experiments in game theory
- KGFR: A Foundation Retriever for Generalized Knowledge Graph Question Answering
- E-CARE: An Efficient LLM-based Commonsense-Augmented Framework for E-Commerce
- An LLM-based Framework for Human-Swarm Teaming Cognition in Disaster Search and Rescue
- BAPPA: Benchmarking Agents, Plans, and Pipelines for Automated Text-to-SQL Generation
- Efficient Reasoning via Thought-Training and Thought-Free Inference
- From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers
- Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
- ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planning
- The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- TabDSR: Decompose, Sanitize, and Reason for Complex Numerical Reasoning in Tabular Data
- Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
- Can LLMs subtract numbers?
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- Random Initialization of Gated Sparse Adapters
- Knowledge Elicitation with Large Language Models for Interpretable Cancer Stage Identification from Pathology Reports
- Aligning LLM agents with human learning and adjustment behavior: a dual agent approach
- ORANGE: An Online Reflection ANd GEneration framework with Domain Knowledge for Text-to-SQL
- How Focused Are LLMs? A Quantitative Study via Repetitive Deterministic Prediction Tasks
- EvoMem: Improving Multi-Agent Planning with Dual-Evolving Memory
- TreeQA: Enhanced LLM-RAG with logic tree reasoning for reliable and interpretable multi-hop question answering
- Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- Diffuse Thinking: Exploring Diffusion Language Models as Efficient Thought Proposers for Reasoning
- Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
- ThoughtProbe: Classifier-Guided LLM Thought Space Exploration via Probing Representations
- A Survey on Generative Recommendation: Data, Model, and Tasks
- Chain of Time: In-Context Physical Simulation with Image Generation Models
- Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
- Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives
- Chain-of-Thought Hijacking
- Pragmatic Theories Enhance Understanding of Implied Meanings in LLMs
- Questionnaire meets LLM: A Benchmark and Empirical Study of Structural Skills for Understanding Questions and Responses
- QuantumBench: A Benchmark for Quantum Problem Solving
- Predicate Renaming via Large Language Models
- Are Language Models Efficient Reasoners? A Perspective from Logic Programming
- TextualVerifier: Verify TextGrad Step-by-Step
- Stemma: Induced Decision Regions Reveal LLM Provenance
- Attention-Aligned Reasoning for Large Language Models
- Two Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local Language Models
- Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
- From Reviews to Actionable Insights: An LLM-Based Approach for Attribute and Feature Extraction
- LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs
- Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks
- Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMs
- Aligning Large Language Models with Procedural Rules: An Autoregressive State-Tracking Prompting for In-Game Trading
- SPICE: Self-Play In Corpus Environments Improves Reasoning
- Quantum Combinatorial Reasoning for Large Language Models
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Uncovering Gaps Between RFC Updates and TCP/IP Implementations: LLM-Facilitated Differential Checks on Intermediate Representations
- Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
- Auto prompting without training labels: An LLM cascade for product quality assessment in e-commerce catalogs
- RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning
- MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal Sarcasm Detection
- Unleashing Diverse Thinking Modes in LLMs through Multi-Agent Collaboration
- Evaluating Large Language Models for Stance Detection on Financial Targets from SEC Filing Reports and Earnings Call Transcripts
- Large language model-based task planning for service robots: A review
- Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
- Can Language Models Compose Skills In-Context?
- Once Upon an Input: Reasoning via Per-Instance Program Synthesis
- Leveraging Large Language Models to Identify Conversation Threads in Collaborative Learning
- MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
- S-Chain: Structured Visual Chain-of-Thought For Medicine
- CHOIR: Collaborative Harmonization fOr Inference Robustness
- Modeling Hierarchical Thinking in Large Reasoning Models
- You Don't Need Prompt Engineering Anymore: The Prompting Inversion
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
- Evaluating Prompting Strategies and Large Language Models in Systematic Literature Review Screening: Relevance and Task-Stage Classification
- How to Auto-optimize Prompts for Domain Tasks? Adaptive Prompting and Reasoning through Evolutionary Domain Knowledge Adaptation
- The Universal Landscape of Human Reasoning
- Reasoning's Razor: Reasoning Improves Accuracy but Can Hurt Recall at Critical Operating Points in Safety and Hallucination Detection
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 Samples
- Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
- Language Ranker: A Lightweight Ranking framework for LLM Decoding
- Code-enabled language models can outperform reasoning models on diverse tasks
- Leveraging the Power of Large Language Models in Entity Linking via Adaptive Routing and Targeted Reasoning
- Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
- Integrating Transparent Models, LLMs, and Practitioner-in-the-Loop: A Case of Nonprofit Program Evaluation
- AgentSense: LLMs Empower Generalizable and Explainable Web-Based Participatory Urban Sensing
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- Timely Clinical Diagnosis through Active Test Selection
- Prompting the Priorities: A First Look at Evaluating LLMs for Vulnerability Triage and Prioritization
- PlanU: Large Language Model Reasoning through Planning under Uncertainty
- DelvePO: Direction-Guided Self-Evolving Framework for Flexible Prompt Optimization
- Rethinking PCA Through Duality
- Language Models as Semantic Augmenters for Sequential Recommenders
- This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch!
- AcademicEval: Live Long-Context LLM Benchmark
- Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Certified Self-Consistency: Statistical Guarantees and Test-Time Training for Reliable Reasoning in LLMs
- Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models
- Investigating Thinking Behaviours of Reasoning-Based Language Models for Social Bias Mitigation
- Physics-Informed Large Language Models for HVAC Anomaly Detection with Autonomous Rule Generation
- Video Reasoning without Training
- Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
- Tutoring LLM into a Better CUDA Optimizer
- T3 Planner: A Self-Correcting LLM Framework for Robotic Motion Planning with Temporal Logic
- DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
- Safe and Efficient In-Context Learning via Risk Control
- Planner and Executor: Collaboration between Discrete Diffusion And Autoregressive Models in Reasoning
- Enhance Large Language Models as Recommendation Systems with Collaborative Filtering
- A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning
- CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
- GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning
- You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
- COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
- Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
- CURE: Confidence-driven Unified Reasoning Ensemble Framework for Medical Question Answering
- Budget-aware Test-time Scaling via Discriminative Verification
- Big Reasoning with Small Models: Instruction Retrieval at Inference Time
- To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Retrieval-in-the-Chain: Bootstrapping Large Language Models for Generative Retrieval
- Adaptive Reasoning Executor: A Collaborative Agent System for Efficient Reasoning
- Schema for In-Context Learning
- Multi-Agent Debate for LLM Judges with Adaptive Stability Detection
- iCodeReviewer: Improving Secure Code Review with Mixture of Prompts
- Self-Verifying Reflection Helps Transformers with CoT Reasoning
- HiCoTraj:Zero-Shot Demographic Reasoning via Hierarchical Chain-of-Thought Prompting from Trajectory
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
- HoneyBee: Data Recipes for Vision-Language Reasoners
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- OneRec-Think: In-Text Reasoning for Generative Recommendation
- ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
- Unlocking the Potential of Diffusion Language Models through Template Infilling
- Automated Skill Decomposition Meets Expert Ontologies: Bridging the Granularity Gap with LLMs
- Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
- Automating Structural Engineering Workflows with Large Language Model Agents
- Evaluating Language Models' Evaluations of Games
- Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
- PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
- Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data
- ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection
- Failure-Driven Workflow Refinement
- Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety
- FINCH: Financial Intelligence using Natural language for Contextualized SQL Handling
- Classifier-Augmented Generation for Structured Workflow Prediction
- LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
- Beyond Surface Reasoning: Unveiling the True Long Chain-of-Thought Capacity of Diffusion Large Language Models
- ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
- Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras
- Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval
- The Idola Tribus of AI: Large Language Models tend to perceive order where none exists
- Verifying Chain-of-Thought Reasoning via Its Computational Graph
- A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages
- Active Model Selection for Large Language Models
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens
- Upfront Chain-of-Thought: A Cooperative Framework for Chain-of-Thought Compression
- GCPO: When Contrast Fails, Go Gold
- ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning
- Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers
- Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
- Guiding Exploration in Reinforcement Learning Through LLM-Augmented Observations
- CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization
- Prepared mind, fast response: A temporal decoupling framework for adaptive knowledge orchestration in open-domain dialogue
- MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
- When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
- Can Speech LLMs Think while Listening?
- MAPRO: Recasting Multi-Agent Prompt Optimization as Maximum a Posteriori Inference
- TS-Agent: Understanding and Reasoning Over Raw Time Series via Iterative Insight Gathering
- EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
- Reasoning for Hierarchical Text Classification: The Case of Patents
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- VelLMes: A high-interaction AI-based deception framework
- Revisiting the Uniform Information Density Hypothesis in LLM Reasoning Traces
- SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
- Unlocking Latent Discourse Translation in LLMs Through Quality-Aware Decoding
- Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
- Do LLMs Know They Are Being Tested? Evaluation Awareness and Incentive-Sensitive Failures in GPT-OSS-20B
- GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
- Adaptive LLM-Symbolic Reasoning via Dynamic Logical Solver Composition
- From Description to Detection: LLM based Extendable O-RAN Compliant Blind DoS Detection in 5G and Beyond
- Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
- Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
- Valid Stopping for LLM Generation via Empirical Dynamic Formal Lift
- LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
- ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems
- GraphGhost: Tracing Structures Behind Large Language Models
- When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
- Code-Switching In-Context Learning for Cross-Lingual Transfer of Large Language Models
- AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering
- SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
- Resource-Efficient Fine-Tuning of LLaMA-3.2-3B for Medical Chain-of-Thought Reasoning
- Natural Language Edge Labelling: Decoupling Intent from Execution in Structured LM Reasoning
- MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts
- Self-Filtered Distillation with LLMs-generated Trust Indicators for Reliable Patent Classification
- Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
- Evaluation of Clinical Trials Reporting Quality using Large Language Models
- QuiLL: An LLM-Based Vulnerability Assessment Framework for the Wild
- From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance
- Beyond Token Length: Step Pruner for Efficient and Accurate Reasoning in Large Language Models
- Generating High-Level Test Cases from Requirements using LLM: An Industry Study
- MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information
- Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods
- GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- Lateral Tree-of-Thoughts Surpasses ToT by Incorporating Logically-Consistent, Low-Utility Candidates
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- NCV: A Node-Wise Consistency Verification Approach for Low-Cost Structured Error Localization in LLM Reasoning
- Knowledge Graph-Guided Multi-Agent Distillation for Reliable Industrial Question Answering with Datasets
- SoT: Structured-of-Thought Prompting Guides Multilingual Reasoning in Large Language Models
- StepChain GraphRAG: Reasoning Over Knowledge Graphs for Multi-Hop Question Answering
- Fine-Tuning on Noisy Instructions: Effects on Generalization and Performance
- GRAD: Generative Retrieval-Aligned Demonstration Sampler for Efficient Few-Shot Reasoning
- PromptPilot: Improving Human-AI Collaboration Through LLM-Enhanced Prompt Engineering
- Rethinking Reward Models for Multi-Domain Test-Time Scaling
- Enhancing Rating Prediction with Off-the-Shelf LLMs Using In-Context User Reviews
- Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generation
- GRPO-λ: Credit Assignment improves LLM Reasoning
- TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
- Adaptive Planning for Multi-Attribute Controllable Summarization with Monte Carlo Tree Search
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- RoBiologyDataChoiceQA: A Romanian Dataset for improving Biology understanding of Large Language Models
- SING-SQL: A Synthetic Data Generation Framework for In-Domain Text-to-SQL Translation
- Nudging the Boundaries of LLM Reasoning
- Unspoken Hints: Accuracy Without Acknowledgement in LLM Reasoning
- Hierarchical Reasoning Models: Perspectives and Misconceptions
- Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
- Plug-and-Play Emotion Graphs for Compositional Prompting in Zero-Shot Speech Emotion Recognition
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Agentic Exploration of Physics Models
- The Dialogue That Heals: A Comprehensive Evaluation of Doctor Agents' Inquiry Capability
- MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language Models
- Learning to Ponder: Adaptive Reasoning in Latent Space
- Advancing mathematics research with generative AI
- Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
- Act as an expert in psychometry. The evaluation of large language models utility in psychological tests cross-cultural adaptations
- ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
- Rethinking and Benchmarking Large Language Models for Graph Reasoning
- ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
- VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
- Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE
- ByteSized32Refactored: Towards an Extensible Interactive Text Games Corpus for LLM World Modeling and Evaluation
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
- Fast Thinking for Large Language Models
- Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
- An Senegalese Legal Texts Structuration Using LLM-augmented Knowledge Graph
- PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
- No Loss, No Gain: Gated Refinement and Adaptive Compression for Prompt Optimization
- MedCritical: Enhancing Medical Reasoning in Small Language Models via Self-Collaborative Correction
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- Tracing Uncertainty in Language Model "Reasoning"
- BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
- UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
- Investigating Faithfulness in Large Audio Language Models
- Green Prompt Engineering: Investigating the Energy Impact of Prompt Design in Software Engineering
- When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
- From Long to Lean: Performance-aware and Adaptive Chain-of-Thought Compression via Multi-round Refinement
- Fuzzy Reasoning Chain (FRC): An Innovative Reasoning Framework from Fuzziness to Clarity
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
- Reimagining Agent-based Modeling with Large Language Model Agents via Shachi
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
- Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
- Why Chain of Thought Fails in Clinical Text Understanding
- Towards Transparent AI: A Survey on Explainable Language Models
- Correct Reasoning Paths Visit Shared Decision Pivots
- Plan2Evolve: LLM Self-Evolution for Improved Planning Capability via Automated Domain Generation
- Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
- Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data
- Bounds of Chain-of-Thought Robustness: Reasoning Steps, Embed Norms, and Beyond
- Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
- Best-of-∞ -- Asymptotic Performance of Test-Time Compute
- UniTransfer: Video Concept Transfer via Progressive Spatial and Timestep Decomposition
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
- Leveraging What's Overfixed: Post-Correction via LLM Grammatical Error Overcorrection
- A Fano-Style Accuracy Upper Bound for LLM Single-Pass Reasoning in Multi-Hop QA
- CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
- Parallel Thinking, Sequential Answering: Bridging NAR and AR for Efficient Reasoning
- Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
- Large Language Models for Real-World IoT Device Identification
- MIXRAG : Mixture-of-Experts Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering
- Thinking While Listening: Simple Test Time Scaling For Audio Classification
- Do Before You Judge: Self-Reference as a Pathway to Better LLM Evaluation
- LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning
- Scaling LLM-Driven Multi-Agent Systems: Design Principles and Architectural Scalability Analysis
- Distilling Answer Set Programming Theories from Large Language Models
- Hierarchical Latent Reasoning for LLM-based Recommendation
- Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
- LASAR: Latent Adaptive Semantic Aligned Reasoning for Generative Recommendation
- Can Large Language Models Resolve Semantic Discrepancy in Self-Destructive Subcultures? Evidence from Jirai Kei
- A Theory of Appropriateness That Accounts for Norms of Rationality
- Language Models Coupled with Metacognition Can Outperform Reasoning Models
- LLMs as verification oracles for Solidity
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- LLM-Enhanced Self-Evolving Reinforcement Learning for Multi-Step E-Commerce Payment Fraud Risk Detection
- Solving Math Word Problems Using Estimation Verification and Equation Generation
- Live-E2T: Real-time Threat Monitoring in Video via Deduplicated Event Reasoning and Chain-of-Thought
- Autonomous Data Agents: A New Opportunity for Smart Data
- Evaluating Large Language Models for Detecting Antisemitism
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
- WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- Automated Knowledge Graph Construction using Large Language Models and Sentence Complexity Modelling
- Program Synthesis via Test-Time Transduction
- Mitigating Strategy-Selection Bias in Reasoning for More Effective Test-Time Scaling
- Adaptive Kernel Design for Bayesian Optimization Is a Piece of CAKE with LLMs
- Automated Facility Enumeration for Building Compliance Checking using Door Detection and Large Language Models
- Adaptive Overclocking: Dynamic Control of Thinking Path Length via Real-Time Reasoning Signals
- Uncovering Implicit Bias in Large Language Models with Concept Learning Dataset
- seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
- KAHAN: Knowledge-Augmented Hierarchical Analysis and Narration for Financial Data Narration
- MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
- Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition-Informed Approach to Quantifying Identity Fusion from Text
- A Novel Differential Feature Learning for Effective Hallucination Detection and Classification
- Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
- Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning
- Towards Universal Debiasing for Language Models-based Tabular Data Generation
- RubikSQL: Lifelong Learning Agentic Knowledge Base as an Industrial NL2SQL System
- Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
- TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies?
- Can LLMs Judge Debates? Evaluating Non-Linear Reasoning via Argumentation Theory Semantics
- DivLogicEval: A Framework for Benchmarking Logical Reasoning Evaluation in Large Language Models
- Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
- MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
- Prompt Stability in Code LLMs: Measuring Sensitivity across Emotion- and Personality-Driven Variations
- Early Stopping Chain-of-thoughts in Large Language Models
- Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
- Improving Context Fidelity via Native Retrieval-Augmented Reasoning
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- REAMS: Reasoning Enhanced Algorithm for Maths Solving
- Evaluating LLM Alignment on Personality Inference from Real-World Interview Data
- The Few-shot Dilemma: Over-prompting Large Language Models
- FedMentor: Domain-Aware Differential Privacy for Heterogeneous Federated LLMs in Mental Health
- Multi-Robot Task Planning for Multi-Object Retrieval Tasks with Distributed On-Site Knowledge via Large Language Models
- Don't Change My View: Ideological Bias Auditing in Large Language Models
- Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
- Root Cause Analysis of Radiation Oncology Incidents Using Large Language Models
- Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
- Pun Unintended: LLMs and the Illusion of Humor Understanding
- GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Models
- POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
- Formal Reasoning for Intelligent QA Systems: A Case Study in the Educational Domain
- LVLMs are Bad at Overhearing Human Referential Communication
- Rethinking Technology Stack Selection with AI Coding Proficiency
- Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
- The System Description of CPS Team for Track on Driving with Language of CVPR 2024 Autonomous Grand Challenge
- LLMAP: LLM-Assisted Multi-Objective Route Planning with User Preferences
- Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
- Unveiling the Latent Directions of Reflection in Large Language Models
- VARCO-VISION-2.0 Technical Report
- Generating Energy-Efficient Code via Large-Language Models -- Where are we now?
- Unsupervised Hallucination Detection by Inspecting Reasoning Processes
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- HieroAction: Hierarchically Guided VLM for Fine-Grained Action Analysis
- XAgents: A Unified Framework for Multi-Agent Cooperation via IF-THEN Rules and Multipolar Task Processing Graph
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- Fluent but Unfeeling: The Emotional Blind Spots of Language Models
- Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference
- TextOnly: A Unified Function Portal for Text-Related Functions on Smartphones
- AI Reasoning for Wireless Communications and Networking: A Survey and Perspectives
- Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals
- Learning from Diverse Reasoning Paths with Routing and Collaboration
- Ensemble Distribution Distillation for Self-Supervised Human Activity Recognition
- SPADE: A Large Language Model Framework for Soil Moisture Pattern Recognition and Anomaly Detection in Precision Agriculture
- Bias after Prompting: Persistent Discrimination in Large Language Models
- Are LLMs Enough for Hyperpartisan, Fake, Polarized and Harmful Content Detection? Evaluating In-Context Learning vs. Fine-Tuning
- PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
- What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects
- How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
- RAFFLES: Reasoning-based Attribution of Faults for LLM Systems
- Can AI Make Energy Retrofit Decisions? An Evaluation of Large Language Models
- From Implicit Exploration to Structured Reasoning: Leveraging Guideline and Refinement for LLMs
- On the Same Wavelength? Evaluating Pragmatic Reasoning in Language Models across Broad Concepts
- Empirical Study of Code Large Language Models for Binary Security Patch Detection
- From Long to Short: LLMs Excel at Trimming Own Reasoning Chains
- DRF: LLM-AGENT Dynamic Reputation Filtering Framework
- Mitigating Spurious Correlations Between Question and Answer via Chain-of-Thought Correctness Perception Distillation
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- ThumbnailTruth: A Multi-Modal LLM Approach for Detecting Misleading YouTube Thumbnails Across Diverse Cultural Settings
- On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
- Chain or tree? Re-evaluating complex reasoning from the perspective of a matrix of thought
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
- ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
- Are We SOLID Yet? An Empirical Study on Prompting LLMs to Detect Design Principle Violations
- DRAssist: Dispute Resolution Assistance using Large Language Models
- HF-RAG: Hierarchical Fusion-based RAG with Multiple Sources and Rankers
- Baichuan-M2: Scaling Medical Capability with Large Verifier System
- Inducing Faithfulness in Structured Reasoning via Counterfactual Sensitivity
- Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
- Rethinking the Chain-of-Thought: The Roles of In-Context Learning and Pre-trained Priors
- MMAPG: A Training-Free Framework for Multimodal Multi-hop Question Answering via Adaptive Planning Graphs
- Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
- LLM-HyPZ: Hardware Vulnerability Discovery using an LLM-Assisted Hybrid Platform for Zero-Shot Knowledge Extraction and Refinement
- RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
- Can Multi-turn Self-refined Single Agent LMs with Retrieval Solve Hard Coding Problems?
- Memory Limitations of Prompt Tuning in Transformers
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database
- Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics
- PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
- How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images
- From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- Re4: Scientific Computing Agent with Rewriting, Resolution, Review and Revision
- MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- Automated Quality Assessment for LLM-Based Complex Qualitative Coding: A Confidence-Diversity Framework
- AI reasoning effort predicts human decision time in content moderation
- AI for Statutory Simplification: A Comprehensive State Legal Corpus and Labor Benchmark
- Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
- Do LVLMs Know What They Know? A Systematic Study of Knowledge Boundary Perception in LVLMs
- Reasoning LLMs in the Medical Domain: A Literature Survey
- Beyond the Textual: Generating Coherent Visual Options for MCQs
- Dynamic Collaboration of Multi-Language Models based on Minimal Complete Semantic Units
- Bias Mitigation Agent: Optimizing Source Selection for Fair and Balanced Knowledge Retrieval
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- STRATA-TS: Selective Knowledge Transfer for Urban Time Series Forecasting with Retrieval-Guided Reasoning
- TrackRec: Iterative Alternating Feedback with Chain-of-Thought via Preference Alignment for Recommendation
- Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation
- Language Models For Generalised PDDL Planning: Synthesising Sound and Programmatic Policies
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- RETAIL: Towards Real-world Travel Planning for Large Language Models
- When Audio and Text Disagree: Revealing Text Bias in Large Audio-Language Models
- Transduction is All You Need for Structured Data Workflows
- Trust but Verify! A Survey on Verification Design for Test-time Scaling
- The Prompting Brain: Neurocognitive Markers of Expertise in Guiding Large Language Models
- Embedding Democratic Values into Social Media AIs via Societal Objective Functions
- Improved Generalized Planning with LLMs through Strategy Refinement and Reflection
- Exit Stories: Using Reddit Self-Disclosures to Understand Disengagement from Problematic Communities
- CausalPlan: Empowering Efficient LLM Multi-Agent Collaboration Through Causality-Driven Planning
- MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers
- G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
- OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
- Can Large Models Teach Student Models to Solve Mathematical Problems Like Human Beings? A Reasoning Distillation Method via Multi-LoRA Interaction
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- TaoSR1: The Thinking Model for E-commerce Relevance Search
- Synchronization Dynamics of Heterogeneous, Collaborative Multi-Agent AI Systems
- Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
- Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow
- Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
- Mitigating Hallucinations in Large Language Models via Causal Reasoning
- RISE: Enhancing VLM Image Annotation with Self-Supervised Reasoning
- Exploring Efficiency Frontiers of Thinking Budget in Medical Reasoning: Scaling Laws between Computational Resources and Reasoning Quality
- Simple o3: Towards Interleaved Vision-Language Reasoning
- Overcoming Knowledge Discrepancies: Structuring Reasoning Threads through Knowledge Balancing in Interactive Scenarios
- Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning
- MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
- HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
- MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
- Ask ChatGPT: Caveats and Mitigations for Individual Users of AI Chatbots
- Benchmark-Driven Selection of AI: Evidence from DeepSeek-R1
- MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
- CS-Agent: LLM-based Community Search via Dual-agent Collaboration
- LLM Empowered Prototype Learning for Zero and Few-Shot Tasks on Tabular Data
- CARES: Collaborative Agentic Reasoning for Error Detection in Surgery
- GreenTEA: Gradient Descent with Topic-modeling and Evolutionary Auto-prompting
- M2LLM: Multi-view Molecular Representation Learning with Large Language Models
- CoDAE: Adapting Large Language Models for Education via Chain-of-Thought Data Augmentation
- \(X\)-evolve: Solution space evolution powered by large language models
- Evaluating Large Language Models as Expert Annotators
- ThinkTuning: Instilling Cognitive Reflections without Distillation
- Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
- CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts
- Keyword-Centric Prompting for One-Shot Event Detection with Self-Generated Rationale Enhancements
- Can You Trick the Grader? Adversarial Persuasion of LLM Judges
- Fine-Tuning Large Language Models Using EEG Microstate Features for Mental Workload Assessment
- Confidence Estimation for Text-to-SQL in Large Language Models
- The benefits and dangers of anthropomorphic conversational agents
- The Problem of Atypicality in LLM-Powered Psychiatry
- Contrastive Analysis of Constituent Order Preferences Within Adverbial Roles in English and Chinese News: A Large-Language-Model-Driven Approach
- UR2: Unify RAG and Reasoning through Reinforcement Learning
- SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs
- PanelTR: Zero-Shot Table Reasoning Framework Through Multi-Agent Scientific Discussion
- Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
- Learning by Teaching: Engaging Students as Instructors of Large Language Models in Computer Science Education
- Omni Geometry Representation Learning vs Large Language Models for Geospatial Entity Resolution
- Large Language Models for Oral History Understanding with Text Classification and Sentiment Analysis
- FineDialFact: A benchmark for Fine-grained Dialogue Fact Verification
- TASE: Token Awareness and Structured Evaluation for Multilingual Language Models
- Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
- ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
- EncQA: Benchmarking Vision-Language Models on Visual Encodings for Charts
- TRAIL: Joint Inference and Refinement of Knowledge Graphs with Large Language Models
- Are Large Language Models Dynamic Treatment Planners? An In Silico Study from a Prior Knowledge Injection Angle
- Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
- StackPilot: Autonomous Function Agents for Scalable and Environment-Free Code Execution
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
- AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization
- Experimental Analysis of Productive Interaction Strategy with ChatGPT: User Study on Function and Project-level Code Generation Tasks
- Enhancing Serendipity Recommendation System by Constructing Dynamic User Knowledge Graphs with Large Language Models
- LLMs are Single-threaded Reasoners: Demystifying the Working Mechanism of Soft Thinking
- Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling
- On the Evaluation of Large Language Models in Multilingual Vulnerability Repair
- Simple Methods Defend RAG Systems Well Against Real-World Attacks
- Everyone Contributes! Incentivizing Strategic Cooperation in Multi-LLM Systems via Sequential Public Goods Games
- Evaluating Position Bias in Large Language Model Recommendations
- From Stimuli to Minds: Enhancing Psychological Reasoning in LLMs via Bilateral Reinforcement Learning
- A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
- TIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language Models
- Accelerating LLM Reasoning via Early Rejection with Partial Reward Modeling
- ProCut: LLM Prompt Compression via Attribution Estimation
- Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
- DeepVIS: Bridging Natural Language and Data Visualization Through Step-wise Reasoning
- How Far Are LLMs from Symbolic Planners? An NLP-Based Perspective
- Better Call Claude: Can LLMs Detect Changes of Writing Style?
- SynAdapt: Learning Adaptive Reasoning in Large Language Models via Synthetic Continuous Chain-of-Thought
- The Missing Parts: Augmenting Fact Verification with Half-Truth Detection
- Thinking Machines: Mathematical Reasoning in the Age of LLMs
- Diagnostic Accuracy of Open-Source Vision-Language Models on Diverse Medical Imaging Tasks
- GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models
- MemoCue: Empowering LLM-Based Agents for Human Memory Recall via Strategy-Guided Querying
- MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models
- Failures Are the Stepping Stones to Success: Enhancing Few-Shot In-Context Learning by Leveraging Negative Samples
- Automatically discovering heuristics in a complex SAT solver with large language models
- An Explainable Emotion Alignment Framework for LLM-Empowered Agent in Metaverse Service Ecosystem
- ChemDFM-R: A Chemical Reasoning LLM Enhanced with Atomized Chemical Knowledge
- Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
- Introducing HALC: A general pipeline for finding optimal prompting strategies for automated coding with LLMs in the computational social sciences
- Can large language models assist choice modelling? Insights into prompting strategies and current models capabilities
Discussions
- Comes from Kojima et al, 2022
arxiv.org/abs/2205.11916 [bsky, 12 points, 1 comments]
- It's effectively a prompting trick - it started out as "think step by step" but now it's baked into the models (with special reinforcement learning to help them come up with better thinking traces bas [bsky, 4 points, 1 comments]
- Large Language Models are Zero-Shot Reasoners: Let's think step by step [hn, 4 points, 0 comments]
- Large Language Models are Zero-Shot Reasoners [lobsters, 4 points, 0 comments]
- ב arxiv.org/pdf/2205.11916 מראים שאפשר פשוט לבקש תעשה דרך. וב arxiv.org/pdf/2310.01714 מראים שאפשר לשפר תוצאות אם מבקשים "להיזכר במקרים דומים ואז לפתור" מה שדי מקרב אותנו לפרדיגמת ה-ReACT של סוכנים. א [bsky, 3 points, 2 comments]
- arxiv.org/abs/2205.11916 [bsky, 2 points, 0 comments]
- IIRC it all started here: arxiv.org/abs/2205.11916 "CoT reasoning" all descended from the observation that putting "let's think step by step" at the end of prompts tended to give much better answers, [bsky, 2 points, 1 comments]
- Large Language Models Are Zero-Shot Reasoners [hn, 1 points, 0 comments]
- [2205.11916] Large Language Models are Zero-Shot Reasoners | arXiv | Cornell University arxiv.org/abs/2205.11916 [bsky, 0 points, 1 comments]
- A veces es cuestión de plantearle al modelo el problema que queremos resolver y decirle "Pensemos este problema paso a paso" para que funcione, como se comenta en este artículo (Kojima et al, 2022).
[bsky, 0 points, 1 comments]
Related