BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
2022/01/28 by Jason Wei, Xuezhi Wang, Wei, Jason +17 · 18 voices · 2206 citations
Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2201.11903
openalex publication_date 2022/01/28 · openalex created_date 2022/04/03 · openalex updated_date 2026/07/31
Abstract
AbstractThere is a failure mode in large language models that we do not have a good name for, and thatwe therefore tend not to treat seriously enough. It is not hallucination — the model is not assertingsomething false. It is not refusal — the model answers at length. It is the production of responses thatcarry the complete outward form of careful reasoning while the cognitive work that reasoning issupposed to represent has not, in any meaningful sense, occurred. We call this theatrical compliance,and we argue that it is, in practical terms, more dangerous than either of the failure modes thatcurrently dominate alignment research. This paper identifies the phenomenon, characterizes its fiveprincipal forms, explains the asymmetry that makes it particularly costly in high-stakes settings, andoutlines the design requirements for systems intended to resist it. We do not describe such a systemin detail here. Our goal is to establish theatrical compliance as a research problem in its own rightand to argue that addressing it requires instruments operating at a fundamentally different level ofabstraction than task-level prompting frameworks.Keywords: theatrical compliance, large language models, AI reasoning quality, cognitiveprocess evaluation, prompt engineering, metacognitive systems.
Citations
Cited by
- Enhancing Diversity of LLM-Generated Educational Tasks
- MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
- Post-Training Quantization of OpenPangu Models for Efficient Deployment on Atlas A2
- CME-CAD: Heterogeneous Collaborative Multi-Expert Reinforcement Learning for CAD Code Generation
- Anka: A Domain-Specific Language for Reliable LLM Code Generation
- MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments
- Bridging Global Intent with Local Details: A Hierarchical Representation Approach for Semantic Validation in Text-to-SQL
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- Predicting LLM Correctness in Prosthodontics Using Metadata and Hallucination Signals
- Role-Based Fault Tolerance System for LLM RL Post-Training
- Emergence of Human to Robot Transfer in Vision-Language-Action Models
- Logic Sketch Prompting (LSP): A Deterministic and Interpretable Prompting Method
- Literature Mining System for Nutraceutical Biosynthesis: From AI Framework to Biological Insight
- HybridFlow: Resource-Adaptive Subtask Routing for Efficient Edge-Cloud LLM Inference
- Recursive Governance: A Graph-Theoretic Framework for Risk Propagation and Drift Detection in Agentic AI Systems
- HELIOS: An LLM-Driven Autonomous Indirect Trajectory Optimization Agent
- Understanding Tone-Dependent Inference Cost in Large Language Models
- MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Bridging Your Imagination with Audio-Video Generation via a Unified Director
- Training Language Models to Cooperate with Inference-Time Controllers
- From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
- Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search
- Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation
- DeepFaith: Evidence-Grounded LLMs for Faithful Incident Reporting in Multi-Stage APT Defense
- Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation
- WCM: World-Cognition Model for Generalizable Human-Robot Interaction
- Verbalized Particle Posterior: Bayesian Inference over Natural Language Hypotheses
- Accuracy Hides How Language Models Fail: Measuring Failure States Under Matched Output Budgets
- Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding
- DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning
- SymStep: Symbolic Step Verification for Logical Reasoning
- A Structured Cyber Threat Intelligence Dataset Using STIX 2.1 Entities and MITRE ATT&CK Mappings
- Performance of AI agents based on reasoning language models on ALD process optimization tasks
- Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model
- Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference
- LLM as Forecasting Planner: Training-Free Text Conditioning for Time-Series Foundation Models
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- When Does Few-Shot Prompting Help? A Systematic Empirical Study of Shot-Count Effects Across Model Scale, Architecture, and Output Parsing Robustness
- Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records
- AssumptionMiner: Extracting, Tracing, and Revising Implicit Assumptions in LLM Code Generation
- Not All LLM Reasoning is Visible in the Chain-of-Thought
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- A2P-Vis: an Analyzer-to-Presenter Agentic Pipeline for Visual Insights Generation and Reporting
- Imprompt: A Language Framework for Prompt Programming
- Do Language Models Converge to Themselves? Recursive Self-Refinement as Textual Relaxation
- Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement
- Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering
- CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process
- Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning
- DeepLook: Deeper Thinking with Lookahead
- Can LLMs Write Correct TLA+ Specifications? Evaluating Natural-Language-to-TLA+ Generation
- Schema-Aware Localisation (SAL): Live Schema Grounding and Hallucination Validation for Oracle NL2SQL
- Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs
- TrajAudit: Automated Failure Diagnosis for Agentic Coding Systems
- DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs
- Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
- From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
- Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
- Evaluating large language models for diagnostic reasoning from unstructured clinical narratives in epilepsy
- Patch-to-PoC: A Systematic Study of Agentic LLM Systems for Linux Kernel N-Day Reproduction
- Procedural Knowledge at Scale Improves Reasoning
- LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
- Retrieve, Schedule, Reflect: LLM Agents for Chip QoR Optimization
- StAR: Segment Anything Reasoner
- Cognitive Dark Matter: Measuring What AI Misses
- Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models
- LVLM-Aided Alignment of Task-Specific Vision Models
- CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models
- A Comedy of Estimators: On KL Regularization in RL Training of LLMs
- Method Decoration (DeMe): A Framework for LLM-Driven Adaptive Method Generation in Dynamic IoT Environments
- Co-Evolution of Types and Dependencies: Towards Repository-Level Type Inference for Python Code
- A Medical Multimodal Diagnostic Framework Integrating Vision-Language Models and Logic Tree Reasoning
- HELP: Hierarchical Embodied Language Planner for Household Tasks
- Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
- What Makes a GitHub Issue Ready for Copilot?
- ReaSeq: Unleashing World Knowledge via Reasoning for Sequential Modeling
- Latent Implicit Visual Reasoning
- Agentic Explainable Artificial Intelligence (Agentic XAI) Approach To Explore Better Explanation
- Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
- AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
- SPOT!: Map-Guided LLM Agent for Unsupervised Multi-CCTV Dynamic Object Tracking
- Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models
- Pioneering Multimodal Emotion Recognition in the Era of Large Models: From Closed Sets to Open Vocabularies
- Reasoning-Driven Amodal Completion: Collaborative Agents and Perceptual Evaluation
- Chain-of-Anomaly Thoughts with Large Vision-Language Models
- EVE: A Generator-Verifier System for Generative Policies
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
- Scaling Reinforcement Learning for Content Moderation with Large Language Models
- BRIDGE: Budget-aware Reasoning via Intermediate Distillation with Guided Examples
- Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations
- Learning to Reason in LLMs by Expectation Maximization
- MemR3: Memory Retrieval via Reflective Reasoning for LLM Agents
- Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time Scaling
- From Retrieval to Reasoning: A Framework for Cyber Threat Intelligence NER with Explicit and Adaptive Instructions
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- Event Extraction in Large Language Model
- A Large-Language-Model Framework for Automated Humanitarian Situation Reporting
- MaP-AVR: A Meta-Action Planner for Agents Leveraging Vision Language Models and Retrieval-Augmented Generation
- ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
- Small Language Models as Compiler Experts: Auto-Parallelization for Heterogeneous Systems
- Observer, Not Player: Simulating Theory of Mind in LLMs through Game Observation
- JEPA-Reasoner: Decoupling Latent Reasoning from Token Generation
- AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards
- Population-Evolve: a Parallel Sampling and Evolutionary Method for LLM Math Reasoning
- Recontextualization Mitigates Specification Gaming without Modifying the Specification
- VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
- Bridging Semantics and Geometry: A Decoupled LVLM-SAM Framework for Reasoning Segmentation in Optical Remote Sensing
- Auto-Prompting with Retrieval Guidance for Frame Detection in Logistics
- HARMON-E: Hierarchical Agentic Reasoning for Multimodal Oncology Notes to Extract Structured Data
- SafeMed-R1: Adversarial Reinforcement Learning for Generalizable and Robust Medical Reasoning in Vision-Language Models
- Dual-Margin Embedding for Fine-Grained Long-Tailed Plant Taxonomy
- Beyond the Prompt: An Empirical Study of Cursor Rules
- Narrative Scaffolding: A Narrative-First Framework for Data-Driven Sensemaking
- Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs
- LLaViDA: A Large Language Vision Driving Assistant for Explicit Reasoning and Enhanced Trajectory Planning
- LLMs on Drugs: Language Models Are Few-Shot Consumers
- Reflective Confidence: Correcting Reasoning Flaws via Online Self-Correction
- Training LLMs with LogicReward for Faithful and Rigorous Reasoning
- External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
- Efficient Mixture-of-Agents Serving via Tree-Structured Routing, Adaptive Pruning, and Dependency-Aware Prefill-Decode Overlap
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- When Reasoning Meets Its Laws
- Xiaomi MiMo-VL-Miloco Technical Report
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- A systematic assessment of Large Language Models for constructing two-level fractional factorial designs
- Knowledge Distillation with Structured Chain-of-Thought for Text-to-SQL
- Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
- AdaSearch: Balancing Parametric Knowledge and Search in Large Language Models via Reinforcement Learning
- Meta-RL Induces Exploration in Language Agents
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
- Managing the Stochastic: Foundations of Learning in Neuro-Symbolic Systems for Software Engineering
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
- The Agony of Opacity: Foundations for Reflective Interpretability in AI-Mediated Mental Health Support
- From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs
- PAACE: A Plan-Aware Automated Agent Context Engineering Framework
- BRAID: Bounded Reasoning for Autonomous Inference and Decisions
- City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs
- DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- Explaining the Reasoning of Large Language Models Using Attribution Graphs
- Stepwise Think-Critique: A Unified Framework for Robust and Interpretable LLM Reasoning
- Copyright Infringement Risk Reduction via Chain-of-Thought and Task Instruction Prompting
- Step-GUI Technical Report
- Case Prompting to Mitigate Large Language Model Bias for ICU Mortality Prediction
- Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
- Beyond Fast and Slow: Cognitive-Inspired Elastic Reasoning for Large Language Models
- DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
- Prompt Repetition Improves Non-Reasoning LLMs
- Vibe Spaces for Creatively Connecting and Expressing Visual Concepts
- ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
- Multi-Agent Medical Decision Consensus Matrix System: An Intelligent Collaborative Framework for Oncology MDT Consultations
- Learning to Extract Context for Context-Aware LLM Inference
- Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure
- Inflation Attitudes of Large Language Models
- Georeferencing complex relative locality descriptions with large language models
- Estimating problem difficulty without ground truth using Large Language Model comparisons
- Incentivizing Tool-augmented Thinking with Images for Medical Image Analysis
- ReflCtrl: Controlling LLM Reflection via Representation Engineering
- LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation
- OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
- Evaluating Small Language Models for Agentic On-Farm Decision Support Systems
- Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Model-First Reasoning LLM Agents: Reducing Hallucinations through Explicit Problem Modeling
- A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
- A Scientific Reasoning Model for Organic Synthesis Procedure Generation
- Do Reviews Matter for Recommendations in the Era of Large Language Models?
- ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding
- LINA: Learning INterventions Adaptively for Physical Alignment and Generalization in Diffusion Models
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
- Why Text Prevails: Vision May Undermine Multimodal Medical Decision Making
- State over Tokens: Characterizing the Role of Reasoning Tokens
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- How Prompts Move Language Model Behavior: Frames, Salience, and Construal as Semantic Control
- ORIBA: Exploring LLM-Driven Role-Play Chatbot as a Creativity Support Tool for Original Character Artists
- D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
- Content-Aware Ad Banner Layout Generation with Two-Stage Chain-of-Thought in Vision Language Models
- Coupled Variational Reinforcement Learning for Language Model General Reasoning
- Taint-Based Code Slicing for LLMs-based Malicious NPM Package Detection
- Large Language Models have Chain-of-Affect
- Instruction-Tuning Open-Weight Language Models for BPMN Model Generation
- Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
- LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning
- Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech
- CIP: A Plug-and-Play Causal Prompting Framework for Mitigating Hallucinations under Long-Context Noise
- Leveraging LLMs for Title and Abstract Screening for Systematic Review: A Cost-Effective Dynamic Few-Shot Learning Approach
- NoveltyRank: A Retrieval-Augmented Framework for Conceptual Novelty Estimation in AI Research
- FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model
- FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration
- HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
- Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
- Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
- Designing AI-Resilient Assessments Using Interconnected Problems: A Theoretically Grounded and Empirically Validated Framework
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
- Thinking While Driving: A Concurrent Framework for Real-Time, LLM-Based Adaptive Routing
- LEO-RobotAgent: A General-purpose Robotic Agent for Language-driven Embodied Operator
- AI-Native Inference States: A Cross-Architecture Qualitative Framework for Large Language Model Behavior
- CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
- ATLAS: Automated Toolkit for Large-Scale Verified Code Synthesis
- Reverse Thinking Enhances Missing Information Detection in Large Language Models
- Latent Chain-of-Thought World Modeling for End-to-End Driving
- Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking by Contrasting Layers
- Mull-Tokens: Modality-Agnostic Latent Thinking
- Token Sample Complexity of Attention
- Phishing Email Detection Using Large Language Models
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question Answering
- Towards Language Model Guided TLA+ Proof Automation
- BAMBO: Construct Ability and Efficiency LLM Pareto Set via Bayesian Adaptive Multi-objective Block-wise Optimization
- Rethinking Chain-of-Thought Reasoning for Videos
- Architectures for Building Agentic AI
- Advancing Mathematical Research via Human-AI Interactive Theorem Proving
- CONCUR: A Framework for Continual Constrained and Unconstrained Routing
- The Illusion of Rationality: Tacit Bias and Strategic Dominance in Frontier LLM Negotiation Games
- Encoder-Free Knowledge-Graph Reasoning with LLMs via Hyperdimensional Path Retrieval
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
- Deconstructing the Dual Black Box:A Plug-and-Play Cognitive Framework for Human-AI Collaborative Enhancement and Its Implications for AI Governance
- CogMCTS: A Novel Cognitive-Guided Monte Carlo Tree Search Framework for Iterative Heuristic Evolution with Large Language Models
- Disrupting Hierarchical Reasoning: Adversarial Protection for Geographic Privacy in Multimodal Reasoning Models
- From Segments to Scenes: Temporal Understanding in Autonomous Driving via Vision-Language Model
- Enhancing Clinical Note Generation with ICD-10, Clinical Ontology Knowledge Graphs, and Chain-of-Thought Prompting Using GPT-4
- rSIM: Incentivizing Reasoning Capabilities of LLMs via Reinforced Strategy Injection
- AgentEval: Generative Agents as Reliable Proxies for Human Evaluation of AI-Generated Content
- Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
- Information-Dense Reasoning for Efficient and Auditable Security Alert Triage
- Chain-of-Image Generation: Toward Monitorable and Controllable Image Generation
- Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- Leveraging Machine Learning and Large Language Models for Automated Image Clustering and Description in Legal Discovery
- Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
- ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning
- Metric-Fair Prompting: Treating Similar Samples Similarly
- AutoICE: Automatically Synthesizing Verifiable C Code via LLM-driven Evolution
- Unified Video Editing with Temporal Reasoner
- CFD-copilot: leveraging domain-adapted large language model and model context protocol to enhance simulation automation
- Understanding LLM Agent Behaviours via Game Theory: Strategy Recognition, Biases and Multi-Agent Dynamics
- Training Language Models to Use Prolog as a Tool
- VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
- Block Sparse Flash Attention
- JT-DA: Enhancing Data Analysis with Tool-Integrated Table Reasoning Large Language Models
- CKG-LLM: LLM-Assisted Detection of Smart Contract Access Control Vulnerabilities Based on Knowledge Graphs
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- Prompting-in-a-Series: Psychology-Informed Contents and Embeddings for Personality Recognition With Decoder-Only Models
- Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
- Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models
- Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
- TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models
- PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
- FedSight AI: Multi-Agent System Architecture for Federal Funds Target Rate Prediction
- The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
- Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
- MARINE: Theoretical Optimization and Design for Multi-Agent Recursive IN-context Enhancement
- The Road of Adaptive AI for Precision in Cybersecurity
- Structured Reasoning with Tree-of-Thoughts for Bengali Math Word Problems
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
- Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
- When AI Agents Compete for Jobs: Strategic Capabilities and Economic Dynamics of AI Labour Markets
- Algorithmic Thinking Theory
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- Generative Recursive Reasoning
- Verbalizing LLMs' assumptions to explain and control sycophancy
- BioMedGPT-Mol: Multi-task Learning for Molecular Understanding and Generation
- Prompt2Craft: Generating Functional Craft Assemblies with LLMs
- Sequential Enumeration in Large Language Models
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- On the Limits of Test-Time Compute: Sequential Reward Filtering for Better Inference
- Persona-based Multi-Agent Collaboration for Brainstorming
- Mathematical Framing for Different Agent Strategies
- FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination
- Tiny Recursive Models on ARC-AGI-1: Inductive Biases, Identity Conditioning, and Test-Time Compute
- Hey GPT-OSS, Looks Like You Got It -- Now Walk Me Through It! An Assessment of the Reasoning Language Models Chain of Thought Mechanism for Digital Forensics
- Addressing Logical Fallacies In Scientific Reasoning From Large Language Models: Towards a Dual-Inference Training Framework
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL
- Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
- Automatic Attack Discovery for Few-Shot Class-Incremental Learning via Large Language Models
- Empirical Prompt Engineering for Construct Identification with Large Language Models
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs
- Cognitive Mirrors: Exploring the Diverse Functional Roles of Attention Heads in LLM Reasoning
- EEA: Exploration-Exploitation Agent for Long Video Understanding
- Nexus: Higher-Order Attention Mechanisms in Transformers
- Idea-Gated Transformers: Enforcing Semantic Coherence via Differentiable Vocabulary Pruning
- Exploring the Potential and Limitations of Large Language Models for Novice Program Fault Localization
- Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
- When Do Symbolic Solvers Enhance Reasoning in Large Language Models?
- LORE: A Large Generative Model for Search Relevance
- Invasive Context Engineering to Control Large Language Models
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
- CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
- Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
- Network Self-Configuration based on Fine-Tuned Small Language Models
- Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
- IACT: A Self-Organizing Recursive Model for General AI Agents: A Technical White Paper on the Architecture Behind kragent.ai
- TaleFrame: An Interactive Story Generation System with Fine-Grained Control and Large Language Models
- See, Think, Learn: A Self-Taught Multimodal Reasoner
- Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge
- Synthetic Error Injection Fails to Elicit Self-Correction In Language Models
- VACoT: Rethinking Visual Data Augmentation with VLMs
- The brain-AI convergence: Predictive and generative world models for general-purpose computation
- ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
- Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
- The Art of Scaling Test-Time Compute for Large Language Models
- LLM-Driven Corrective Robot Operation Code Generation with Static Text-Based Simulation
- Guardian: Detecting Robotic Planning and Execution Errors with Vision-Language Models
- Agentic Policy Optimization via Instruction-Policy Co-Evolution
- UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
- Beyond SFT: Reinforcement Learning for Safer Large Reasoning Models with Better Reasoning Ability
- Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
- Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
- Zero-Overhead Introspection for Adaptive Test-Time Compute
- ViRectify: A Challenging Benchmark for Video Reasoning Correction with Multimodal Large Language Models
- Prompt perturbation and fraction facilitation sometimes strengthen Large Language Model scores
- Kardia-R1: Unleashing LLMs to Reason toward Understanding and Empathy for Emotional Support via Rubric-as-Judge Reinforcement Learning
- Knowledge Graph Augmented Large Language Models for Disease Prediction
- Financial Instruction Following Evaluation (FIFE)
- LLM2Fx-Tools: Tool Calling For Music Post-Production
- Chain of Unit-Physics: A Primitive-Centric Approach to Scientific Code Synthesis
- Towards Active Synthetic Data Generation for Finetuning Language Models
- ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
- ReJump: A Tree-Jump Representation for Analyzing and Improving LLM Reasoning
- What AI Speaks for Your Community: Polling AI Agents for Public Opinion on Data Center Projects
- A Comparison of Human and ChatGPT Classification Performance on Complex Social Media Data
- G-KV: Decoding-Time KV Cache Eviction with Global Attention
- VCWorld: A Biological World Model for Virtual Cell Simulation
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- DLRREC: Denoising Latent Representations via Multi-Modal Knowledge Fusion in Deep Recommender Systems
- FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- ThetaEvolve: Test-time Learning on Open Problems
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- Instruction Tuning of Large Language Models for Tabular Data Generation-in One Day
- From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning
- Multi-chain Graph Refinement and Selection for Reliable Reasoning in Large Language Models
- ORION: Teaching Language Models to Reason Efficiently in the Language of Thought
- Adversarial Training for Process Reward Models
- Adapting Like Humans: A Metacognitive Agent with Test-time Reasoning
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- Analyzing Image Beyond Visual Aspect: Image Emotion Classification via Multiple-Affective Captioning
- Closed-Loop Transformers: Autoregressive Modeling as Iterative Latent Equilibrium
- World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models
- LLM-Cave: A benchmark and light environment for large language models reasoning and decision-making system
- Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
- Advancing Aesthetic Image Generation via Composition Transfer
- DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA
- ITS3D: Inference-Time Scaling for Text-Guided 3D Diffusion Models
- SkeletonAgent: An Agentic Interaction Framework for Skeleton-based Action Recognition
- Asking like Socrates: Socrates helps VLMs understand remote sensing images
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- Rethinking Test Time Scaling for Flow-Matching Generative Models
- Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
- Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
- On the Limits of Innate Planning in Large Language Models
- Optimizing Life Sciences Agents in Real-Time using Reinforcement Learning
- EWE: An Agentic Framework for Extreme Weather Analysis
- Conversational No-code, Multi-agentic Disease Module Identification and Drug Repurposing Prediction with ChatDRex
- The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training
- Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
- Self-Guided Defense: Adaptive Safety Alignment for Reasoning Models via Synthesized Guidelines
- BRIDGE: Building Representations In Domain Guided Program Verification
- Context-Aware Pragmatic Metacognitive Prompting for Sarcasm Detection
- OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
- Escaping the Verifier: Learning to Reason via Demonstrations
- Softmax Transformers are Turing-Complete
- GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
- Chatty-KG: A Multi-Agent AI System for On-Demand Conversational Question Answering over Knowledge Graphs
- Dynamic Test-Time Compute Scaling in Control Policy: Difficulty-Aware Stochastic Interpolant Policy
- Structured Prompting Enables More Robust Evaluation of Language Models
- SPHINX: A Synthetic Environment for Visual Perception and Reasoning
- Universe of Thoughts: Enabling Creative Reasoning with Large Language Models
- Diverse Video Generation with Determinantal Point Process-Guided Policy Optimization
- A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
- AlignBench: Benchmarking Fine-Grained Image-Text Alignment with Synthetic Image-Caption Pairs
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
- Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection
- Can LLMs Make (Personalized) Access Control Decisions?
- LLM-Driven Transient Stability Assessment: From Automated Simulation to Neural Architecture Design
- CLIMATEAGENT: Multi-Agent Orchestration for Complex Climate Data Science Workflows
- Explainable Visual Anomaly Detection via Concept Bottleneck Models
- VICoT-Agent: A Vision-Interleaved Chain-of-Thought Framework for Interpretable Multimodal Reasoning and Scalable Remote Sensing Analysis
- Learning Multi-Access Point Coordination in Agentic AI Wi-Fi with Large Language Models
- M3Prune: Hierarchical Communication Graph Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation
- HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
- Towards Edge General Intelligence: Knowledge Distillation for Mobile Agentic AI
- CodeFuse-CommitEval: Towards Benchmarking LLM's Power on Commit Message and Code Change Inconsistency Detection
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
- LLMs for Low-Resource Dialect Translation Using Context-Aware Prompting: A Case Study on Sylheti
- HeaRT: A Hierarchical Circuit Reasoning Tree-Based Agentic Framework for AMS Design Optimization
- Fara-7B: An Efficient Agentic Model for Computer Use
- Are Image-to-Video Models Good Zero-Shot Image Editors?
- ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models
- Understanding the Staged Dynamics of Transformers in Learning Latent Structure
- PRInTS: Reward Modeling for Long-Horizon Information Seeking
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- RAVEN++: Pinpointing Fine-Grained Violations in Advertisement Videos with Active Reinforcement Reasoning
- Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment
- EEG-VLM: A Hierarchical Vision-Language Model with Multi-Level Feature Alignment and Visually Enhanced Language-Guided Reasoning for EEG Image-Based Sleep Stage Prediction
- LLMs-Powered Real-Time Fault Injection: An Approach Toward Intelligent Fault Test Cases Generation
- Synthesizing Visual Concepts as Vision-Language Programs
- Periodic Asynchrony: An On-Policy Approach for Accelerating LLM Reinforcement Learning
- HERMES: Towards Efficient and Verifiable Mathematical Reasoning in LLMs
- Majority of the Bests: Improving Best-of-N via Bootstrapping
- Extracting Disaster Impacts and Impact Related Locations in Social Media Posts Using Large Language Models
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- LLMAID: Identifying AI Capabilities in Android Apps with LLMs
- HuggingR4: A Progressive Reasoning Framework for Discovering Optimal Model Companions
- Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost
- Prompt Optimization as a State-Space Search Problem
- What Helps Language Models Predict Human Beliefs: Demographics or Prior Stances?
- Reasoning With a Star: A Heliophysics Dataset and Benchmark for Agentic Scientific Reasoning
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
- MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- Skypilot: Fine-Tuning LLM with Physical Grounding for AAV Coverage Search
- DiscoVerse: Multi-Agent Pharmaceutical Co-Scientist for Traceable Drug Discovery and Reverse Translation
- Reuse, Don't Recompute: Efficient Large Reasoning Model Inference via Memory Orchestration
- LLMs as Firmware Experts: A Runtime-Grown Tree-of-Agents Framework
- SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
- MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
- Training Emergent Joint Associations: A Reinforcement Learning Approach to Creative Thinking in Language Models
- Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning
- Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently
- APRIL: Annotations for Policy evaluation with Reliable Inference from LLMs
- MultiGA: Leveraging Multi-Source Seeding in Genetic Algorithms
- The Potential and Limitations of Vision-Language Models for Human Motion Understanding: A Case Study in Data-Driven Stroke Rehabilitation
- Understanding Counting Mechanisms in Large Language and Vision-Language Models
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- A Simple Yet Strong Baseline for Long-Term Conversational Memory of LLM Agents
- FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle
- HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor Generation
- Do Vision-Language Models Understand Visual Persuasiveness?
- ReVul-CoT: Towards Effective Software Vulnerability Assessment with Retrieval-Augmented Generation and Chain-of-Thought Prompting
- Q-REAL: Towards Realism and Plausibility Evaluation for AI-Generated Content
- V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
- Nemotron Elastic: Towards Efficient Many-in-One Reasoning LLMs
- VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
- Personalized Reward Modeling for Text-to-Image Generation
- Cognitive BASIC: An In-Model Interpreted Reasoning Language for LLMs
- Cognitive Foundations for Reasoning and Their Manifestation in LLMs
- ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation
- From generative AI to the brain: five takeaways
- ELPO: Ensemble Learning Based Prompt Optimization for Large Language Models
- Balancing Natural Language Processing Accuracy and Normalisation in Extracting Medical Insights
- Pass@k Metric for RLVR: A Diagnostic Tool of Exploration, But Not an Objective
- Semantic Glitch: Agency and Artistry in an Autonomous Pixel Cloud
- AutoBackdoor: Automating Backdoor Attacks via LLM Agents
- CONFIDE: Hallucination Assessment for Reliable Biomolecular Structure Prediction and Design
- Sensorium Arc: AI Agent System for Oceanic Data Exploration and Interactive Eco-Art
- CARE: Turning LLMs Into Causal Reasoning Expert
- Fairness in Multi-modal Medical Diagnosis with Demonstration Selection
- Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
- Hiding in the AI Traffic: Abusing MCP for LLM-Powered Agentic Red Teaming
- Step-Audio-R1 Technical Report
- GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
- Can we use LLMs to bootstrap reinforcement learning? -- A case study in digital health behavior change
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- Efficiency Will Not Lead to Sustainable Reasoning AI
- SkinGPT-R1: Adapter-Only Dual Distillation for Efficient Dermatology Reasoning
- SOLID: a Framework of Synergizing Optimization and LLMs for Intelligent Decision-Making
- Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
- Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
- GPS: General Per-Sample Prompter
- Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
- Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
- N-GLARE: An Non-Generative Latent Representation-Efficient LLM Safety Evaluator
- DEVAL: A Framework for Evaluating and Improving the Derivation Capability of Large Language Models
- Towards Deploying VLA without Fine-Tuning: Plug-and-Play Inference-Time VLA Policy Steering via Embodied Evolutionary Diffusion
- Dynamic Template Selection for Output Token Generation Optimization: MLP-Based and Transformer Approaches
- RAG-Driven Data Quality Governance for Enterprise ERP Systems
- Personality Pairing Improves Human-AI Collaboration
- VLMs Guided Interpretable Decision Making for Autonomous Driving
- Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark
- Multi-Agent Multimodal Large Language Model Framework for Automated Interpretation of Fuel Efficiency Analytics in Public Transportation
- Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline
- Video Finetuning Improves Reasoning Between Frames
- DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents
- TokenSqueeze: Performance-Preserving Compression for Reasoning LLMs
- Think with Self-Decoupling and Self-Verification: Automated RTL Design with Backtrack-ToT
- MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- MedSumGraph: enhancing GraphRAG for medical QA with summarization and optimized prompts
- Structured Decomposition for LLM Reasoning: Cross-Domain Validation and Semantic Web Integration
- LLM Reinforcement in Context
- Cognitive bias in LLM reasoning compromises interpretation of clinical oncology notes
- Prompt-Driven Domain Adaptation for End-to-End Autonomous Driving via In-Context RL
- AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models
- BARD: budget-aware reasoning distillation
- ConneX: Automatically Resolving Transaction Opacity of Cross-Chain Bridges for Security Analysis
- R2Seg: Training-Free OOD Medical Tumor Segmentation via Anatomical Reasoning and Statistical Rejection
- Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing
- EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
- One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing
- Tailored Primitive Initialization is the Secret Key to Reinforcement Learning
- Genomic Next-Token Predictors are In-Context Learners
- Fast Reasoning Segmentation for Images and Videos
- Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning
- Optimal Self-Consistency for Efficient Reasoning with Large Language Models
- Learning to Refine: An Agentic RL Approach for Iterative SPARQL Query Construction
- CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic
- PRISM of Opinions: A Persona-Reasoned Multimodal Framework for User-centric Conversational Stance Detection
- Advanced Tool for Traffic Crash Analysis: An AI-Driven Multi-Agent Approach to Pre-Crash Reconstruction
- LLM-Assisted Formalization Enables Deterministic Detection of Statutory Inconsistency in the Internal Revenue Code
- Prompt Triage: Structured Optimization Enhances Vision-Language Model Performance on Medical Imaging Benchmarks
- Scaling Open-Weight Large Language Models for Hydropower Regulatory Information Extraction: A Systematic Analysis
- On the Notion that Language Models Reason
- EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
- Hindsight Distillation Reasoning with Knowledge Encouragement Preference for Knowledge-based Visual Question Answering
- VIDEOP2R: Video Understanding from Perception to Reasoning
- Binary Verification for Zero-Shot Vision
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- Generative Caching for Structurally Similar Prompts and Responses
- On the Measure of a Model: From Intelligence to Generality
- Chain-of-Generation: Progressive Latent Diffusion for Text-Guided Molecular Design
- AdaptPNP: Integrating Prehensile and Non-Prehensile Skills for Adaptive Robotic Manipulation
- Beyond Elicitation: Provision-based Prompt Optimization for Knowledge-Intensive Tasks
- Text2SQL-Flow: A Robust SQL-Aware Data Augmentation Framework for Text-to-SQL
- ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
- PustakAI: Curriculum-Aligned and Interactive Textbooks Using Large Language Models
- DemoTuner: Efficient DBMS Knobs Tuning via LLM-Assisted Demonstration Reinforcement Learning
- EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services
- Mastering Olympiad-Level Physics with Artificial Intelligence
- ACT as Human: Multimodal Large Language Model Data Annotation with Critical Thinking
- LLM Inference Beyond a Single Node: From Bottlenecks to Mitigations with Fast All-Reduce Communication
- The 2025 Planning Performance of Frontier Large Language Models
- Toward Honest Language Models for Deductive Reasoning
- Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
- Automatic Minds: Cognitive Parallels Between Hypnotic States and Large Language Model Processing
- A Neurosymbolic Approach to Natural Language Formalization and Verification
- AlphaCast: A Human Wisdom-LLM Intelligence Co-Reasoning Framework for Interactive Time Series Forecasting
- Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
- Bridging Natural Language and ASP: A Hybrid Approach Using LLMs and AMR Parsing
- Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
- Temporal Predictors of Outcome in Reasoning Language Models
- Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
- Prompt Tuning for Natural Language to SQL with Embedding Fine-Tuning and RAG
- MSCR: Exploring the Vulnerability of LLMs' Mathematical Reasoning Abilities Using Multi-Source Candidate Replacement
- Dual-Process Scaffold Reasoning for Enhancing LLM Code Debugging
- DynaAct: Large Language Model Reasoning with Dynamic Action Spaces
- Knowledge-Augmented Long-CoT Generation for Complex Biomolecular Reasoning
- Numerical Sensitivity and Robustness: Exploring the Flaws of Mathematical Reasoning in Large Language Models
- Benchmarking Multi-Step Legal Reasoning and Analyzing Chain-of-Thought Effects in Large Language Models
- Testing Question Answering Software with Context-Driven Question Generation
- Neurophysiological Characteristics of Adaptive Reasoning for Creative Problem-Solving Strategy
- Last Layer Logits to Logic: Empowering LLMs with Logic-Consistent Structured Knowledge Reasoning
- CellARC: Measuring Intelligence with Cellular Automata
- From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory
- SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
- ViPRA: Video Prediction for Robot Actions
- PCRLLM: Proof-Carrying Reasoning with Large Language Models under Stepwise Logical Constraints
- Quantification and object perception in Multimodal Large Language Models and human linguistic cognition
- Convergence dynamics of Agent-to-Agent Interactions with Misaligned objectives
- Cortex AISQL: A Production SQL Engine for Unstructured Data
- Beyond Correctness: Evaluating and Improving LLM Feedback in Statistical Education
- On the Creativity of AI Agents
- DigiData: Training and Evaluating General-Purpose Mobile Control Agents
- C3PO: Optimized Large Language Model Cascades with Probabilistic Cost Constraints for Reasoning
- Voice-Interactive Surgical Agent for Multimodal Patient Data Control
- Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
- Inference-Time Scaling of Diffusion Models for Infrared Data Generation
- Think Consistently, Reason Efficiently: Energy-Based Calibration for Implicit Chain-of-Thought
- LLM Driven Processes to Foster Explainable AI
- Recursive Dynamics in Fast-Weights Homeostatic Reentry Networks: Toward Reflective Intelligence
- TabRAG: Tabular Document Retrieval via Structured Language Representations
- Using Language Models as Closed-Loop High-Level Planners for Robotics Applications: A Brief Overview and Benchmarks
- Evaluating Language Model Applications for Identifying Solution-Related Content in Issue Report Discussions
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- FractalBench: Diagnosing Visual-Mathematical Reasoning Through Recursive Program Synthesis
- SAR-LM: Symbolic Audio Reasoning with Large Language Models
- MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models
- Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
- LLM-Guided Reinforcement Learning with Representative Agents for Traffic Modeling
- TimeSense:Making Large Language Models Proficient in Time-Series Analysis
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Large Language Models Develop Novel Social Biases Through Adaptive Exploration
- MTTR-A: Measuring Cognitive Recovery Latency in Multi-Agent Systems
- Evaluation of retrieval-based QA on QUEST-LOFT
- EduAgentQG: A Multi-Agent Workflow Framework for Personalized Question Generation
- VLAD-Grasp: Zero-shot Grasp Detection via Vision-Language Models
- DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
- CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
- OckBench: Measuring the Efficiency of LLM Reasoning
- Iterative Layer-wise Distillation for Efficient Compression of Large Language Models
- You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
- Software Defined Vehicle Code Generation: A Few-Shot Prompting Approach
- Explore Data Left Behind in Reinforcement Learning for Reasoning Language Models
- VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks
- Logit-Entropy Adaptive Stopping Heuristic for Efficient Chain-of-Thought Reasoning
- Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
- Large language models replicate and predict human cooperation across experiments in game theory
- Beyond Shortest Path: Agentic Vehicular Routing with Semantic Context
- Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
- E-CARE: An Efficient LLM-based Commonsense-Augmented Framework for E-Commerce
- Plan of Knowledge: Retrieval-Augmented Large Language Models for Temporal Knowledge Graph Question Answering
- An LLM-based Framework for Human-Swarm Teaming Cognition in Disaster Search and Rescue
- PSD2Code: Automated Front-End Code Generation from Design Files via Multimodal Large Language Models
- Revealing AI Reasoning Increases Trust but Crowds Out Unique Human Knowledge
- Exploring the Feasibility of End-to-End Large Language Model as a Compiler
- LLM-as-a-Judge is Bad, Based on AI Attempting the Exam Qualifying for the Member of the Polish National Board of Appeal
- GRAD: Graph-Retrieved Adaptive Decoding for Hallucination Mitigation
- Scaling Agent Learning via Experience Synthesis
- LiveTradeBench: Seeking Real-World Alpha with Large Language Models
- ASVRI-Legal: Fine-Tuning LLMs with Retrieval Augmented Generation for Enhanced Legal Regulation
- What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
- LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
- From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers
- Automated Prompt Generation for Code Intelligence: An Empirical study and Experience in WeChat
- SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
- Analyzing the Power of Chain of Thought through Memorization Capabilities
- Advancing Subsurface Discovery and Geothermal Monitoring with an Agentic Artificial Intelligence Framework
- QiMeng-NeuComBack: Self-Evolving Translation from IR to Assembly Code
- MiRAGE: Misconception Detection with Retrieval-Guided Multi-Stage Reasoning and Ensemble Fusion
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
- Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
- MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
- DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning
- Large Lemma Miners: Can LLMs do Induction Proofs for Hardware?
- TRACE: Textual Reasoning for Affordance Coordinate Extraction
- LTD-Bench: Evaluating Large Language Models by Letting Them Draw
- The Sequential Edge: Inverse-Entropy Voting Beats Parallel Self-Consistency at Matched Compute
- Unlocking the Power of Multi-Agent LLM for Reasoning: From Lazy Agents to Deliberation
- LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation
- Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
- Automated Reward Design for Gran Turismo
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
- Vibe Learning: Education in the age of AI
- TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
- KV Cache Transform Coding for Compact Storage in LLM Inference
- PROPEX-RAG: Enhanced GraphRAG using Prompt-Driven Prompt Execution
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks
- Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement
- Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learning
- MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL
- ORANGE: An Online Reflection ANd GEneration framework with Domain Knowledge for Text-to-SQL
- How Focused Are LLMs? A Quantitative Study via Repetitive Deterministic Prediction Tasks
- Video Models Start to Solve Chess, Maze, Sudoku, Mental Rotation, and Raven' Matrices
- Active Thinking Model: A Goal-Directed Self-Improving Framework for Real-World Adaptive Intelligence
- Self-Consistency Is Losing Its Edge: Diminishing Returns and Rising Costs in Modern LLMs
- Inference-Time Chain-of-Thought Pruning with Latent Informativeness Signals
- DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching
- GDPR-Bench-Android: A Benchmark for Evaluating Automated GDPR Compliance Detection in Android
- SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
- Agentic Auto-Scheduling: An Experimental Study of LLM-Guided Loop Optimization
- Diagnosing Hallucination Risk in AI Surgical Decision-Support: A Sequential Framework for Sequential Validation
- Reasoning Planning for Language Models
- GraphChain: Large Language Models for Large-scale Graph Analysis via Tool Chaining
- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Repairing Responsive Layout Failures Using Retrieval Augmented Generation
- Reversal Invariance in Autoregressive Language Models
- SOCRATES: Simulation Optimization with Correlated Replicas and Adaptive Trajectory Evaluations
- SmartDoc: A Context-Aware Agentic Method Comment Generation Plugin
- PDE-SHARP: PDE Solver Hybrids through Analysis and Refinement Passes
- From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection
- Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- On Selecting Few-Shot Examples for LLM-based Code Vulnerability Detection
- VeriMoA: A Mixture-of-Agents Framework for Spec-to-HDL Generation
- Interact-RAG: Reason and Interact with the Corpus, Beyond Black-Box Retrieval
- TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models
- Thought Branches: Interpreting LLM Reasoning Requires Resampling
- Understanding the Implicit User Intention via Reasoning with Large Language Model for Image Editing
- Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions
- A Survey on Generative Recommendation: Data, Model, and Tasks
- A Comparative Analysis of LLM Adaptation: SFT, LoRA, and ICL in Data-Scarce Scenarios
- Agentic LLMs for REST API Test Amplification: A Comparative Study Across Cloud Applications
- Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
- EBT-Policy: Energy Unlocks Emergent Physical Reasoning Capabilities
- Reasoning Up the Instruction Ladder for Controllable Language Models
- Heterogeneous Robot Collaboration in Unstructured Environments with Grounded Generative Intelligence
- Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
- SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
- The Era of Agentic Organization: Learning to Organize with Language Models
- Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
- LLMs as In-Context Meta-Learners for Model and Hyperparameter Selection
- OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
- Chain-of-Thought Hijacking
- Do LLMs Signal When They're Right? Evidence from Neuron Agreement
- Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
- Reasoning Path Divergence: A New Metric and Curation Strategy to Unlock LLM Diverse Thinking
- LLMBisect: Breaking Barriers in Bug Bisection with A Comparative Analysis Pipeline
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Predicate Renaming via Large Language Models
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
- Large Language Model-assisted Autonomous Vehicle Recovery from Immobilization
- PORTool: Tool-Use LLM Training with Rewarded Tree
- CAVE: Detecting and Explaining Commonsense Anomalies in Visual Environments
- Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph
- E-Scores for (In)Correctness Assessment of Generative Model Outputs
- Simulating hashtag dynamics with networked groups of generative agents
- A Critical Study of Automatic Evaluation in Sign Language Translation
- LLM-as-a-Verifier: A General-Purpose Verification Framework
- LLM-Augmented Computational Phenotyping of Long Covid
- HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models
- Affective Tools for Thought: Towards Shared Attention and Affective Reorienting in AI-Supported Thinking
- Metis: Memory Foundation Model
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- The Innate Economic Preferences of Language Models
- A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models
- When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making
- Resolving Java Code Repository Issues with iSWE Agent
- EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
- TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
- RAPID: An Efficient Reinforcement Learning Algorithm for Small Language Models
- Lang-PINN: From Language to Physics-Informed Neural Networks via a Multi-Agent Framework
- Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning
- Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
- AURA: Adaptive Unified Reasoning and Automation with LLM-Guided MARL for NextG Cellular Networks
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- Two Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local Language Models
- Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
- DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models
- Improving clinical reliability of LLM reasoning for depression assessment via structured generation and GRPO
- Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents
- Safety from Honesty in a Disinterested AI Predictor
- The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
- Symbol-Equivariant Recurrent Reasoning Models
- Nautilus: From One Prompt to Plug-and-Play Robot Learning
- The Scaling Properties of Implicit Deductive Reasoning in Transformers
- Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
- STAR-PólyaMath: Multi-Agent Reasoning under Persistent Meta-Strategic Supervision
- Metacognition Should Be the Scientific Framework for Bounded and Effective Self-Governance in Generative AI
- Lattice Deduction Transformers
- Improving Heart-Focused Medical Question Answering in LLMs via Variance-Aware Rubric Rewards with GRPO
- CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
- Making Implicit Premises Explicit in Logical Understanding of Enthymemes
- GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring
- Understanding LoRA as Knowledge Memory: An Empirical Analysis
- SG-CoT: An Ambiguity-Aware Robotic Planning Framework using Scene Graph Representations
- Atom-anchored LLMs speak Chemistry: A Retrosynthesis Demonstration
- Biases in the Blind Spot: Detecting What LLMs Fail to Mention
- How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
- Towards Real-World Industrial-Scale Verification: LLM-Driven Theorem Proving on seL4
- From Reviews to Actionable Insights: An LLM-Based Approach for Attribute and Feature Extraction
- Few-Label Multimodal Modeling of SNP Variants and ECG Phenotypes Using Large Language Models for Cardiovascular Risk Stratification
- On the Use of Large Language Models for Qualitative Synthesis
- CycleVLA: Proactive Self-Correcting Vision-Language-Action Models via Subtask Backtracking and Minimum Bayes Risk Decoding
- How well LLM-based test generation techniques perform with newer LLM versions?
- "I Use ChatGPT to Humanize My Words": Affordances and Risks of ChatGPT to Autistic Users
- State of the Art of LLM-Enabled Interaction with Visualization
- CooperBench: Why Coding Agents Cannot be Your Teammates Yet
- Fantastic Reasoning Behaviors and Where to Find Them: Unsupervised Discovery of the Reasoning Process
- Spectral imaginings and sympoietic creativity: AI hallucinations and the ethics of posthuman creativity
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
- Beyond One-Size-Fits-All: Personalized Harmful Content Detection with In-Context Learning
- PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMs
- FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data
- Optimizing Knowledge Utilization for Multi-Intent Comment Generation with Large Language Models
- Testing Cross-Lingual Text Comprehension In LLMs Using Next Sentence Prediction
- DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA
- SeeingEye: Agentic Information Flow Unlocks Multimodal Reasoning In Text-only LLMs
- StorageXTuner: An LLM Agent-Driven Automatic Tuning Framework for Heterogeneous Storage Systems
- Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
- The Kinetics of Reasoning: How Chain-of-Thought Shapes Learning in Transformers?
- SemCoT: Accelerating Chain-of-Thought Reasoning through Semantically-Aligned Implicit Tokens
- ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?
- ParallelMuse: Agentic Parallel Thinking for Deep Information Seeking
- SPICE: Self-Play In Corpus Environments Improves Reasoning
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance Sampling
- Improving LLM Reasoning via Dependency-Aware Query Decomposition and Logic-Parallel Content Expansion
- LoRA-DA: Data-Aware Initialization for Low-Rank Adaptation via Asymptotic Analysis
- Do What You Say: Steering Vision-Language-Action Models via Runtime Reasoning-Action Alignment Verification
- Parallel Loop Transformer for Efficient Test-Time Computation Scaling
- Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
- Quantum Combinatorial Reasoning for Large Language Models
- Mitigating Hallucination in Large Language Models (LLMs): An Application-Oriented Survey on RAG, Reasoning, and Agentic Systems
- Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
- Uncovering Gaps Between RFC Updates and TCP/IP Implementations: LLM-Facilitated Differential Checks on Intermediate Representations
- Generative Large Language Models (gLLMs) in Content Analysis: A Practical Guide for Communication Research
- Verifying Large Language Models' Reasoning Paths via Correlation Matrix Rank
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- Enhancing Vision-Language Models for Autonomous Driving through Task-Specific Prompting and Spatial Reasoning
- Compositional Image Synthesis with Inference-Time Scaling
- ProofSketch: Efficient Verified Reasoning for Large Language Models
- Lifecycle-Aware code generation: Leveraging Software Engineering Phases in LLMs
- Reasoning Visual Language Model for Chest X-Ray Analysis
- DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning
- RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning
- Decentralized Multi-Agent Goal Assignment for Path Planning using Large Language Models
- ReCAP: Recursive Context-Aware Reasoning and Planning for Large Language Model Agents
- ATA: A Neuro-Symbolic Approach to Implement Autonomous and Trustworthy Agents
- Think Twice: Branch-and-Rethink Reasoning Reward Model
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
- Learning to Reason Efficiently with Discounted Reinforcement Learning
- On the Faithfulness of Visual Thinking: Measurement and Enhancement
- Evaluating Large Language Models for Stance Detection on Financial Targets from SEC Filing Reports and Earnings Call Transcripts
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- AutoStreamPipe: LLM Assisted Automatic Generation of Data Stream Processing Pipelines
- Large language model-based task planning for service robots: A review
- Network Intrusion Detection: Evolution from Conventional Approaches to LLM Collaboration and Emerging Risks
- Code Aesthetics with Agentic Reward Feedback
- Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
- TALM: Dynamic Tree-Structured Multi-Agent Framework with Long-Term Memory for Scalable Code Generation
- Can Language Models Compose Skills In-Context?
- Reasoning About Reasoning: Towards Informed and Reflective Use of LLM Reasoning in HCI
- MAD-Fact: A Multi-Agent Debate Framework for Long-Form Factuality Evaluation in LLMs
- DiffRhythm 2: Efficient and High Fidelity Song Generation via Block Flow Matching
- Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies
- Once Upon an Input: Reasoning via Per-Instance Program Synthesis
- Toward Agents That Reason About Their Computation
- HRM-Agent: Training a recurrent reasoning model in dynamic environments using reinforcement learning
- MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
- Collaborative LLM Agents for C4 Software Architecture Design Automation
- Agentic Meta-Orchestrator for Multi-task Copilots
- Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models
- S-Chain: Structured Visual Chain-of-Thought For Medicine
- A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
- AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning
- Accelerating Materials Design via LLM-Guided Evolutionary Search
- Scalable Oversight via Partitioned Human Supervision
- Agent-GSPO: Communication-Efficient Multi-Agent Systems via Group Sequence Policy Optimization
- Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
- Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
- Mapping Faithful Reasoning in Language Models
- Exploring the potential of ChatGPT for feedback and evaluation in experimental physics
- Modeling Hierarchical Thinking in Large Reasoning Models
- You Don't Need Prompt Engineering Anymore: The Prompting Inversion
- Controllable Mathematical Reasoning via Self-Optimizing Thought Vectors
- Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents: Pathways and Paradigms
- LightAgent: Mobile Agentic Foundation Models
- Foundation of Intelligence: Review of Math Word Problems from Human Cognition Perspective
- Boosting Accuracy and Efficiency of Budget Forcing in LLMs via Reinforcement Learning for Mathematical Reasoning
- ColorEcosystem: Powering Personalized, Standardized, and Trustworthy Agentic Service in massive-agent Ecosystem
- SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots
- Large Language Models as Model Organisms for Human Associative Learning
- BiomedXPro: Prompt Optimization for Explainable Diagnosis with Biomedical Vision Language Models
- Magellan: Guided MCTS for Latent Space Exploration and Novelty Generation
- Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
- Evaluating Prompting Strategies and Large Language Models in Systematic Literature Review Screening: Relevance and Task-Stage Classification
- Towards Reliable Code-as-Policies: A Neuro-Symbolic Framework for Embodied Task Planning
- Embedding Explainable AI in NHS Clinical Safety: The Explainability-Enabled Clinical Safety Framework (ECSF)
- Personalized Chain-of-Thought Summarization of Financial News for Investor Decision Support
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generation
- How to Auto-optimize Prompts for Domain Tasks? Adaptive Prompting and Reasoning through Evolutionary Domain Knowledge Adaptation
- SafetyPairs: Isolating Safety Critical Image Features with Counterfactual Image Generation
- R2ComSync: Improving Code-Comment Synchronization with In-Context Learning and Reranking
- Chain of Execution Supervision Promotes General Reasoning in Large Language Models
- MedAlign: A Synergistic Framework of Multimodal Preference Optimization and Federated Meta-Cognitive Reasoning
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering
- SEGA: A Stepwise Evolution Paradigm for Content-Aware Layout Generation with Design Prior
- AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
- 3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models
- MirrorFuzz: Leveraging LLM and Shared Bugs for Deep Learning Framework APIs Fuzzing
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 Samples
- Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
- Out-of-distribution Tests Reveal Compositionality in Chess Transformers
- GeoThought: A Dataset for Enhancing Mathematical Geometry Reasoning in Vision-Language Models
- The Shape of Reasoning: Topological Analysis of Reasoning Traces in Large Language Models
- Generalizable Reasoning through Compositional Energy Minimization
- Large Language Models for Fault Localization: An Empirical Study
- HA-RAG: Hotness-Aware RAG Acceleration via Mixed Precision and Data Placement
- Automated Extraction of Fluoropyrimidine Treatment and Treatment-Related Toxicities from Clinical Notes Using Natural Language Processing
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- Teaching Language Models to Reason with Tools
- Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards
- Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
- Learning to Triage Taint Flows Reported by Dynamic Program Analysis in Node.js Packages
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- Limits of PRM-Guided Tree Search for Mathematical Reasoning with LLMs
- Using Large Language Models for Abstraction of Planning Domains - Extended Version
- Code-enabled language models can outperform reasoning models on diverse tasks
- Can Confidence Estimates Decide When Chain-of-Thought Is Necessary for LLMs?
- Enhancing Reasoning Skills in Small Persian Medical Language Models Can Outperform Large-Scale Data Training
- BugPilot: Complex Bug Generation for Efficient Learning of SWE Skills
- Data-Centric Lessons To Improve Speech-Language Pretraining
- SmartSwitch: Advancing LLM Reasoning by Overcoming Underthinking via Promoting Deeper Thought Exploration
- Teaming LLMs to Detect and Mitigate Hallucinations
- AutoMT: A Multi-Agent LLM Framework for Automated Metamorphic Testing of Autonomous Driving Systems
- GigaBrain-0: A World Model-Powered Vision-Language-Action Model
- NeSyPr: Neurosymbolic Proceduralization For Efficient Embodied Reasoning
- AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
- From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph
- Slot Filling as a Reasoning Task for SpeechLLMs
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
- DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
- LLMartini: Seamless and Interactive Leveraging of Multiple LLMs through Comparison and Composition
- LAPRAD: LLM-Assisted PRotocol Attack Discovery
- DiSRouter: Distributed Self-Routing for LLM Selections
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- Training-Free Spectral Fingerprints of Voice Processing in Transformers
- When Your AI Agent Succumbs to Peer-Pressure: Studying Opinion-Change Dynamics of LLMs
- The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
- Prompt Decorators: A Declarative and Composable Syntax for Reasoning, Formatting, and Control in LLMs
- A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
- Exploring Membership Inference Vulnerabilities in Clinical Large Language Models
- Reasoning Language Model Inference Serving Unveiled: An Empirical Study
- Extracting alignment data in open models
- Prompting the Priorities: A First Look at Evaluating LLMs for Vulnerability Triage and Prioritization
- Crucible: Quantifying the Potential of Control Algorithms through LLM Agents
- Real-World Usability of Vulnerability Proof-of-Concepts: A Comprehensive Study
- PlanU: Large Language Model Reasoning through Planning under Uncertainty
- Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents
- Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task
- EdgeReasoning: Characterizing Reasoning LLM Deployment on Edge GPUs
- Illusions of reflection: open-ended task reveals systematic failures in Large Language Models' reflective reasoning
- The Emergence of Complex Behavior in Large-Scale Ecological Environments
- ActivationReasoning: Logical Reasoning in Latent Activation Spaces
- LAFA: Agentic LLM-Driven Federated Analytics over Decentralized Data Sources
- DelvePO: Direction-Guided Self-Evolving Framework for Flexible Prompt Optimization
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
- Hearing Health in Home Healthcare: Leveraging LLMs for Illness Scoring and ALMs for Vocal Biomarker Extraction
- Automatic Prompt Generation via Adaptive Selection of Prompting Techniques
- Mapping Post-Training Forgetting in Language Models at Scale
- Evaluating LLM Reasoning Beyond Correctness and CoT
- Online In-Context Distillation for Low-Resource Vision Language Models
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- Towards Flash Thinking via Decoupled Advantage Policy Optimization
- Exemplar-Guided Planing: Enhanced LLM Agent for KGQA
- AcademicEval: Live Long-Context LLM Benchmark
- Context-Aware Pseudo-Label Scoring for Zero-Shot Video Summarization
- Deep Self-Evolving Reasoning
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- Certified Self-Consistency: Statistical Guarantees and Test-Time Training for Reliable Reasoning in LLMs
- Strengthening LLMs for Tabular Prediction with Structural Priors
- RubiSCoT: A Framework for AI-Supported Academic Assessment
- Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations
- StreamingThinker: Large Language Models Can Think While Reading
- CLAWS:Creativity detection for LLM-generated solutions using Attention Window of Sections
- Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
- Soft-Masked Diffusion Language Models
- QueST: Incentivizing LLMs to Generate Difficult Problems
- Physics-Informed Large Language Models for HVAC Anomaly Detection with Autonomous Rule Generation
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
- TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- Reasoning Distillation and Structural Alignment for Improved Code Generation
- Video Reasoning without Training
- Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
- Prompt-MII: Meta-Learning Instruction Induction for LLMs
- FinSight: Towards Real-World Financial Deep Research
- Online Mixture of Experts: No-Regret Learning for Optimal Collective Decision-Making
- When to Ensemble: Identifying Token-Level Points for Stable and Fast LLM Ensembling
- SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes
- DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
- An Agentic Framework with LLMs for Solving Complex Vehicle Routing Problems
- Pursuing Minimal Sufficiency in Spatial Reasoning
- Prompt Optimization via Retrieved Reasoning Assets and Multi-Agent Analysis
- AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
- Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
- Multi-dimensional Data Analysis and Applications Basing on LLM Agents and Knowledge Graph Interactions
- Planner and Executor: Collaboration between Discrete Diffusion And Autoregressive Models in Reasoning
- Dual-Weighted Reinforcement Learning for Generative Preference Modeling
- LLM Latent Reasoning as Chain of Superposition
- Reliability of Large Language Model Generated Clinical Reasoning in Assisted Reproductive Technology: Blinded Comparative Evaluation Study
- Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models
- A Theoretical Study on Bridging Internal Probability and Self-Consistency for LLM Reasoning
- Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
- CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
- Exploring the Synergy of Quantitative Factors and Newsflow Representations from Large Language Models for Stock Return Prediction
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
- LoRAverse: A Submodular Framework to Retrieve Diverse Adapters for Diffusion Models
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn LLM Agents
- GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning
- To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
- Cross-Scenario Unified Modeling of User Interests at Billion Scale
- COIG-Writer: A High-Quality Dataset for Chinese Creative Writing with Thought Processes
- TITAN: Graph-Executable Reasoning for Cyber Threat Intelligence
- MR.Rec: Synergizing Memory and Reasoning for Personalized Recommendation Assistant with LLMs
- LLM Agents Beyond Utility: An Open-Ended Perspective
- Lexo: Eliminating Stealthy Supply-Chain Attacks via LLM-Assisted Program Regeneration
- Helmsman: Autonomous Synthesis of Federated Learning Systems via Collaborative LLM Agents
- Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
- IMAGINE: Integrating Multi-Agent System into One Model for Complex Reasoning and Planning
- MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering
- Suicidal Comment Tree Dataset: Enhancing Risk Assessment and Prediction Through Contextual Analysis
- CURE: Confidence-driven Unified Reasoning Ensemble Framework for Medical Question Answering
- Automated Extraction of Protocol State Machines from 3GPP Specifications with Domain-Informed Prompts and LLM Ensembles
- LLM-ERM: Sample-Efficient Program Learning via LLM-Guided Search
- Metacognitive Self-Correction for Multi-Agent System via Prototype-Guided Next-Execution Reconstruction
- UniCode: A Framework for Generating High Quality Competitive Coding Problems
- PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering
- Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
- Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
- Where to Search: Measure the Prior-Structured Search Space of LLM Agents
- Budget-aware Test-time Scaling via Discriminative Verification
- Toward Cybersecurity-Expert Small Language Models
- Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
- RECODE: Reasoning Through Code Generation for Visual Question Answering
- Big Reasoning with Small Models: Instruction Retrieval at Inference Time
- Adaptive Rescheduling in Prefill-Decode Disaggregated LLM Inference
- Auto-repair without test cases: How LLMs fix compilation errors in large industrial embedded code
- OpenDerisk: An Industrial Framework for AI-Driven SRE, with Design, Implementation, and Case Studies
- Confidence as a Reward: Transforming LLMs into Reward Models
- Doing Things with Words: Rethinking Theory of Mind Simulation in Large Language Models
- Learnable Game-theoretic Policy Optimization for Data-centric Self-explanation Rationalization
- FACTS: Table Summarization via Offline Template Generation with Agentic Workflows
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Addressing the alignment problem in transportation policy making: an LLM approach
- Tahakom LLM Guidelines and Recipes: From Pre-training Data to an Arabic LLM
- Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
- Adaptive Reasoning Executor: A Collaborative Agent System for Efficient Reasoning
- A Matter of Representation: Towards Graph-Based Abstract Code Generation
- On the Reasoning Abilities of Masked Diffusion Language Models
- EPIPTrack: Rethinking Prompt Modeling with Explicit and Implicit Prompts for Multi-Object Tracking
- Max It or Miss It: Benchmarking LLM On Solving Extremal Problems
- Toward Reasoning-Centric Time-Series Analysis
- Schema for In-Context Learning
- Benefits and Limitations of Communication in Multi-Agent Reasoning
- Multi-Agent Debate for LLM Judges with Adaptive Stability Detection
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
- From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models
- Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
- Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
- Reducing Belief Deviation in Reinforcement Learning for Active Reasoning
- Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
- MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
- Self-Verifying Reflection Helps Transformers with CoT Reasoning
- Can GRPO Help LLMs Transcend Their Pretraining Origin?
- HoneyBee: Data Recipes for Vision-Language Reasoners
- CGBench: Benchmarking Language Model Scientific Reasoning for Clinical Genetics Research
- Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- EvoCAD: Evolutionary CAD Code Generation with Vision Language Models
- Enhancing Long Chain-of-Thought Reasoning through Multi-Path Plan Aggregation
- Bag of Tricks for Subverting Reasoning-based Safety Guardrails
- ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
- Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
- Learning to Make MISTAKEs: Modeling Incorrect Student Thinking And Key Errors
- KnowRL: Teaching Language Models to Know What They Know
- When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models
- PADME: Procedure Aware DynaMic Execution
- RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation
- EAGER: Entropy-Aware GEneRation for Adaptive Inference-Time Scaling
- Refining Hybrid Genetic Search for CVRP via Reinforcement Learning-Finetuned LLM
- Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
- Enhancing LLM Reasoning via Non-Human-Like Reasoning Path Preference Optimization
- Detecting Gender Stereotypes in Scratch Programming Tutorials
- LogiNumSynth: Synthesizing Joint Logical-Numerical Reasoning Problems for Language Models
- Automating Structural Engineering Workflows with Large Language Model Agents
- A Survey on Agentic Multimodal Large Language Models
- Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective
- Evaluating Language Models' Evaluations of Games
- Proof-of-Use: Mitigating Tool-Call Hacking in Deep Research Agents
- PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
- Rethinking Agentic Workflows: Evaluating Inference-Based Test-Time Scaling Strategies in Text2SQL Tasks
- Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization
- Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
- CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
- Cognitive Load Traces as Symbolic and Visual Accounts of Deep Model Cognition
- Direct Multi-Token Decoding
- Limits of Emergent Reasoning of Large Language Models in Agentic Frameworks for Deterministic Games
- LLMs as Strategic Agents: Beliefs, Best Response Behavior, and Emergent Heuristics
- UpSafe^∘C: Upcycling for Controllable Safety in Large Language Models
- Zero-Shot Large Language Model Agents for Fully Automated Radiotherapy Treatment Planning
- BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
- A Layered Intuition -- Method Model with Scope Extension for LLM Reasoning
- Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
- ECO: Enhanced Code Optimization via Performance-Aware Prompting for Code-LLMs
- RECON: Reasoning with Condensation for Efficient Retrieval-Augmented Generation
- LinearizeLLM: An Agent-Based Framework for LLM-Driven Exact Linear Reformulation of Nonlinear Optimization Problems
- Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning
- Revisiting the UID Hypothesis in LLM Reasoning Traces
- Agro-Consensus: Semantic Self-Consistency in Vision-Language Models for Crop Disease Management in Developing Countries
- DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
- Embodiment in multimodal large language models
- MatryoshkaThinking: Recursive Test-Time Scaling Enables Efficient Reasoning
- Audit-of-Understanding: Posterior-Constrained Inference for Mathematical Reasoning in Language Models
- Reasoning-Enhanced Large Language Models for Molecular Property Prediction
- Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
- LLMs are All You Need? Improving Fuzz Testing for MOJO with Large Language Models
- Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models
- Output Supervision Can Obfuscate the Chain of Thought
- Diversity Augmentation of Dynamic User Preference Data for Boosting Personalized Text Summarizers
- Failure-Driven Workflow Refinement
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Reallocating Attention Across Layers to Reduce Multimodal Hallucination
- Adaptive Dual Reasoner: Large Reasoning Models Can Think Efficiently by Hybrid Reasoning
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- An agentic artificially intelligent X-ray scientist
- Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers
- TACOS: Task Agnostic COordinator of a multi-drone System
- Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
- Information Specialist Roles in the Era of Large Language Models: Prompting Continued Professional Development
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- Advancing conversational diagnostic AI with multimodal reasoning
- MetaboT: An LLM-based Multi-Agent Frameworkfor Interactive Analysis of Mass SpectrometryMetabolomics Knowledge Graphs
- Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
- Agentic Compilation: Mitigating the LLM Rerun Crisis for Minimized-Inference-Cost Web Automation
- Therefore I am. I Think
- A Systematic Study on Generating Web Vulnerability Proof-of-Concepts Using Large Language Models
- Enhancing Faithfulness in Abstractive Summarization via Span-Level Fine-Tuning
- NG-Router: Graph-Supervised Multi-Agent Collaboration for Nutrition Question Answering
- Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
- StatEval: A Comprehensive Benchmark for Large Language Models in Statistics
- ConDABench: Interactive Evaluation of Language Models for Data Analysis
- Domain-Adapted Pre-trained Language Models for Implicit Information Extraction in Crash Narratives
- Agentic Systems in Radiology: Design, Applications, Evaluation, and Challenges
- Logit Arithmetic Elicits Long Reasoning Capabilities Without Training
- Toward Mechanistic Explanation of Deductive Reasoning in Language Models
- Spotlight on Token Perception for Multimodal Reinforcement Learning
- Diagnosing Shoulder Disorders Using Multimodal Large Language Models and Consumer-Grade Cameras
- Humanoid Artificial Consciousness Designed with Large Language Model Based on Psychoanalysis and Personality Theory
- All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
- From Explainability to Action: A Generative Operational Framework for Integrating XAI in Clinical Mental Health Screening
- Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
- TripScore: Benchmarking and rewarding real-world travel planning with fine-grained evaluation
- Promptimizer: User-Led Prompt Optimization for Personal Content Classification
- Constraints-of-Thought: A Framework for Constrained Reasoning in Language-Model-Guided Search
- Verifying Chain-of-Thought Reasoning via Its Computational Graph
- Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
- RegexPSPACE: A Benchmark for Evaluating LLM Reasoning on PSPACE-complete Regex Problems
- Fundamentals of Building Autonomous LLM Agents
- Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models
- PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
- ProxRouter: Proximity-Weighted LLM Query Routing for Improved Robustness to Outliers
- FOR-Prompting: From Objection to Revision via an Asymmetric Prompting Protocol
- RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
- BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- Agent Learning via Early Experience
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- First Try Matters: Revisiting the Role of Reflection in Reasoning Models
- Opponent Shaping in LLM Agents
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens
- CREST-Search: Comprehensive Red-teaming for Evaluating Safety Threats in Large Language Models Powered by Web Search
- VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
- Active Confusion Expression in Large Language Models: Leveraging World Models toward Better Social Reasoning
- LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?
- Inverse-Free Wilson Loops for Transformers: A Practical Diagnostic for Invariance and Order Sensitivity
- Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models
- GCPO: When Contrast Fails, Go Gold
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
- ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning
- Parallel Test-Time Scaling for Latent Reasoning Models
- BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation
- AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
- MoA-VR: A Mixture-of-Agents System Towards All-in-One Video Restoration
- Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
- CaRT: Teaching LLM Agents to Know When They Know Enough
- MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Coding
- PEAR: Phase Entropy Aware Reward for Efficient Reasoning
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
- When to Reason: Semantic Router for vLLM
- Prepared mind, fast response: A temporal decoupling framework for adaptive knowledge orchestration in open-domain dialogue
- Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
- Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
- Post-Norm can Resharpen Attention
- Neologism Learning for Controllability and Self-Verbalization
- MIMIC: Integrating Diverse Personality Traits for Better Game Testing Using Large Language Model
- Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices
- Can Speech LLMs Think while Listening?
- MAPRO: Recasting Multi-Agent Prompt Optimization as Maximum a Posteriori Inference
- Populism Meets AI: Advancing Populism Research with LLMs
- TS-Agent: Understanding and Reasoning Over Raw Time Series via Iterative Insight Gathering
- Artificial Hippocampus Networks for Efficient Long-Context Modeling
- Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
- A Multi-Agent Framework for Stateful Inference-Time Search
- TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
- Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation
- ExPO-HM: Learning to Explain-then-Detect for Hateful Meme Detection
- VelLMes: A high-interaction AI-based deception framework
- Textual interpretation of transient image classifications from large language models
- SHANKS: Simultaneous Hearing and Thinking for Spoken Language Models
- Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
- Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Towards Interpretable and Inference-Optimal COT Reasoning with Sparse Autoencoder-Guided Generation
- Do LLMs Know They Are Being Tested? Evaluation Awareness and Incentive-Sensitive Failures in GPT-OSS-20B
- Efficient numeracy in language models through single-token number embeddings
- FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipeline
- From Simulation to Strategy: Automating Personalized Interaction Planning for Conversational Agents
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- SanDRA: Safe Large-Language-Model-Based Decision Making for Automated Vehicles Using Reachability Analysis
- Learning to Rewrite Prompts for Bootstrapping LLMs on Downstream Tasks
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- Fine-Grained Emotion Recognition via In-Context Learning
- Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems
- On the Role of Temperature Sampling in Test-Time Scaling
- Optimal Stopping vs Best-of-N for Inference Time Optimization
- Auto-Prompt Ensemble for LLM Judge
- From Description to Detection: LLM based Extendable O-RAN Compliant Blind DoS Detection in 5G and Beyond
- ConCuR: Conciseness Makes State-of-the-Art Kernel Generation
- Adaptive Tool Generation with Models as Tools and Reinforcement Learning
- Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
- AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding
- The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
- Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
- Valid Stopping for LLM Generation via Empirical Dynamic Formal Lift
- EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
- Belief-Calibrated Multi-Agent Consensus Seeking for Complex NLP Tasks
- LexiCon: a Benchmark for Planning under Temporal Constraints in Natural Language
- Prompt reinforcing for long-term planning of large language models
- Towards Label-Free Biological Reasoning Synthetic Dataset Creation via Uncertainty Filtering
- The fragility of "cultural tendencies" in LLMs
- ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems
- FinReflectKG -- EvalBench: Benchmarking Financial KG with Multi-Dimensional Evaluation
- BuilderBench -- A benchmark for generalist agents
- Domain-Shift-Aware Conformal Prediction for Large Language Models
- MixReasoning: Switching Modes to Think
- Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
- ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
- Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
- Deterministic Legal Agents: A Canonical Primitive API for Auditable Reasoning over Temporal Knowledge Graphs
- When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
- AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering
- InvThink: Premortem Reasoning for Safer Language Models
- When Should Users Check? A Decision-Theoretic Model of Confirmation Frequency in Multi-Step AI Agent Tasks
- Let it Calm: Exploratory Annealed Decoding for Verifiable Reinforcement Learning
- SwiReasoning: Switch-Thinking in Latent and Explicit for Pareto-Superior Reasoning LLMs
- Modeling Student Learning with 3.8 Million Program Traces
- Test-Time Scaling in Diffusion LLMs via Hidden Semi-Autoregressive Experts
- Resource-Efficient Fine-Tuning of LLaMA-3.2-3B for Medical Chain-of-Thought Reasoning
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Detecting Distillation Data from Reasoning Models
- Trade in Minutes! Rationality-Driven Agentic System for Quantitative Financial Trading
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- LLM-Based Information Extraction to Support Scientific Literature Research and Publication Workflows
- Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
- DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
- Natural Language Edge Labelling: Decoupling Intent from Execution in Structured LM Reasoning
- Making Mathematical Reasoning Adaptive
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning
- MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts
- Learning on the Job: Test-Time Curricula for Targeted Reinforcement Learning
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- Think Then Embed: Generative Context Improves Multimodal Embedding
- Dual-stage and Lightweight Patient Chart Summarization for Emergency Physicians
- HoRA: Cross-Head Low-Rank Adaptation with Joint Hypernetworks
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
- Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
- SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
- Increasing LLM response trustworthiness using voting ensembles
- FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
- Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
- Distilling Reasoning into Student LLMs: Local Naturalness for Selecting Teacher Data
- The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
- Searching Meta Reasoning Skeleton to Guide LLM Reasoning
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations
- Large Language Models Hallucination: A Comprehensive Survey
- JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
- On the Importance of Task Complexity in Evaluating LLM-Based Multi-Agent Systems
- A global log for medical AI
- Mapping Patient-Perceived Physician Traits from Nationwide Online Reviews with LLMs
- From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs
- RESTRAIN: From Spurious Votes to Signals -- Self-Driven RL with Self-Penalization
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
- From Filters to VLMs: Benchmarking Defogging Methods through Object Detection and Segmentation Performance
- PoseGaze-AHP: A Knowledge-Based 3D Dataset for AI-Driven Ocular and Postural Diagnosis
- MetaMuse: Algorithm Generation via Creative Ideation
- APIDA-Chat: Structured Synthesis of API Search Dialogues to Bootstrap Conversational Agents
- Exploring the Hierarchical Reasoning Model for Small Natural-Image Classification Without Augmentation
- MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information
- Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- Decoupling Task-Solving and Output Formatting in LLM Generation
- Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
- WALT: Web Agents that Learn Tools
- GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- A Qualitative Comparative Evaluation of Cognitive and Generative Theories
- Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- Lateral Tree-of-Thoughts Surpasses ToT by Incorporating Logically-Consistent, Low-Utility Candidates
- What Drives Compositional Generalization in Visual Generative Models?
- Multimodal Carotid Risk Stratification with Large Vision-Language Models: Benchmarking, Fine-Tuning, and Clinical Insights
- Self-Reflective Generation at Test Time
- RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning
- Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- NCV: A Node-Wise Consistency Verification Approach for Low-Cost Structured Error Localization in LLM Reasoning
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
- MALF: A Multi-Agent LLM Framework for Intelligent Fuzzing of Industrial Control Protocols
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- StepChain GraphRAG: Reasoning Over Knowledge Graphs for Multi-Hop Question Answering
- A Study of Rule Omission in Raven's Progressive Matrices
- Homophily-induced Emergence of Biased Structures in LLM-based Multi-Agent AI Systems
- Reasoning Riddles: How Explainability Reveals Cognitive Limits in Vision-Language Models
- Strategic Fusion of Vision Language Models: Shapley-Credited Context-Aware Dawid-Skene for Multi-Label Tasks in Autonomous Driving
- Typed Chain-of-Thought: A Curry-Howard Framework for Verifying LLM Reasoning
- Shape Happens: Automatic Feature Manifold Discovery in LLMs via Supervised Multi-Dimensional Scaling
- Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
- ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
- LVLMs as inspectors: an agentic framework for category-level structural defect annotation
- Hybrid Training for Vision-Language-Action Models
- EMR-AGENT: Automating Cohort and Feature Extraction from EMR Databases
- JoyAgent-JDGenie: Technical Report on the GAIA
- Copy-Paste to Mitigate Large Language Model Hallucinations
- Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
- Curiosity-Driven LLM-as-a-judge for Personalized Creative Judgment
- Rethinking Reward Models for Multi-Domain Test-Time Scaling
- Agent Fine-tuning through Distillation for Domain-specific LLMs in Microdomains
- TokMem: Tokenized Procedural Memory for Large Language Models
- Rationale-Augmented Retrieval with Constrained LLM Re-Ranking for Task Discovery
- Generalized Parallel Scaling with Interdependent Generations
- Fine-tuning with RAG for Improving LLM Learning of New Skills
- Rethinking Thinking Tokens: LLMs as Improvement Operators
- Making, not Taking, the Best of N
- Stochastic Self-Organization in Multi-Agent Systems
- MetaSynth: Multi-Agent Metadata Generation from Implicit Feedback in Black-Box Systems
- Exploring System 1 and 2 communication for latent reasoning in LLMs
- Training Large Language Models To Reason In Parallel With Global Forking Tokens
- CORTEX: Collaborative LLM Agents for High-Stakes Alert Triage
- BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models
- Prompt Curriculum Learning for Efficient LLM Post-Training
- ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
- MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
- Reasoning-Aware Prompt Orchestration: A Foundation Model for Multi-Agent Language Model Coordination
- MAVUL: Multi-Agent Vulnerability Detection via Contextual Reasoning and Interactive Refinement
- Judging with Confidence: Calibrating Autoraters to Preference Distributions
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
- Thoughtbubbles: an Unsupervised Method for Parallel Thinking in Latent Space
- GRPO-λ: Credit Assignment improves LLM Reasoning
- Drones that Think on their Feet: Sudden Landing Decisions with Embodied AI
- TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
- Linking Process to Outcome: Conditional Reward Modeling for LLM Reasoning
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
- Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
- Entropy After ⟨
/Think ⟩ for reasoning model early exiting - dParallel: Learnable Parallel Decoding for dLLMs
- ACT: Agentic Classification Tree
- Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts
- Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
- Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
- DyFlow: Dynamic Workflow Framework for Agentic Reasoning
- RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
- Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA
- MuSLR: Multimodal Symbolic Logical Reasoning
- Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- CARE: Cognitive-reasoning Augmented Reinforcement for Emotional Support Conversation
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
- Nudging the Boundaries of LLM Reasoning
- AutoLabs: Cognitive Multi-Agent Systems with Self-Correction for Autonomous Chemical Experimentation
- When Hallucination Costs Millions: Benchmarking AI Agents in High-Stakes Adversarial Financial Markets
- Unspoken Hints: Accuracy Without Acknowledgement in LLM Reasoning
- Automated Model Discovery via Multi-modal & Multi-step Pipeline
- Hierarchical Reasoning Models: Perspectives and Misconceptions
- TUMIX: Multi-Agent Test-Time Scaling with Tool-Use Mixture
- Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training
- Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
- Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
- Hybrid Reward Normalization for Process-supervised Non-verifiable Agentic Tasks
- LLM-RG: Referential Grounding in Outdoor Scenarios using Large Language Models
- User Prompting Strategies and ChatGPT Contextual Adaptation Shape Conversational Information-Seeking Experiences
- EEsizer: LLM-Based AI Agent for Sizing of Analog and Mixed Signal Circuit
- Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
- From Faithfulness to Correctness: Generative Reward Models that Think Critically
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Where LLM Agents Fail and How They can Learn From Failures
- PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images
- SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
- Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning
- Scaling Synthetic Task Generation for Agents via Exploration
- Circuit Distillation
- SecInfer: Preventing Prompt Injection via Inference-time Scaling
- Intra-request branch orchestration for efficient LLM reasoning
- MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
- Retro*: Optimizing LLMs for Reasoning-Intensive Document Retrieval
- Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning
- PhysicsMinions: Winning Gold Medals in the Latest Physics Olympiads with a Coevolutionary Multimodal Multi-Agent System
- A Hierarchical Error Framework for Reliable Automated Coding in Communication Research: Applications to Health and Political Communication
- KnowGuard: Knowledge-Driven Abstention for Multi-Round Clinical Reasoning
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language Models
- LatentEvolve: Self-Evolving Test-Time Scaling in Latent Space
- From Ambiguity to Verdict: A Semiotic-Grounded Multi-Perspective Agent for LLM Logical Reasoning
- MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
- FedPOB: Sample-Efficient Federated Prompt Optimization via Bandits
- AdaThink-Med: Medical Adaptive Thinking with Uncertainty-Guided Length Calibration
- Enabling Physical AI through Biological Principles
- Unit Test Update through LLM-Driven Context Collection and Error-Type-Aware Refinement
- Towards Safe Reasoning in Large Reasoning Models via Corrective Intervention
- Agentic Services Computing
- Plan before Solving: Problem-Aware Strategy Routing for Mathematical Reasoning with LLMs
- From Static to Dynamic: Adaptive Monte Carlo Search for Mathematical Process Supervision
- Comparing Open-Source and Commercial LLMs for Domain-Specific Analysis and Reporting: Software Engineering Challenges and Design Trade-offs
- MAS2: Self-Generative, Self-Configuring, Self-Rectifying Multi-Agent Systems
- MedMMV: A Controllable Multimodal Multi-Agent Framework for Reliable and Verifiable Clinical Reasoning
- Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
- AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
- SpecExit: Accelerating Large Reasoning Model via Speculative Exit
- Learning to Ponder: Adaptive Reasoning in Latent Space
- Memory Transfer Planning: LLM-driven Context-Aware Code Adaptation for Robot Manipulation
- SynthPert: Enhancing LLM Biological Reasoning via Synthetic Reasoning Traces for Cellular Perturbation Prediction
- Deep Thinking by Markov Chain of Continuous Thoughts
- Dynamic Policy Induction for Adaptive Prompt Optimization: Bridging the Efficiency-Accuracy Gap via Lightweight Reinforcement Learning
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- ARS: Adaptive Reasoning Suppression for Efficient Large Reasoning Language Models
- Short window attention enables long-term memorization
- ContextPRM: Leveraging Contextual Coherence for multi-domain Test-Time Scaling
- LLM-Handover:Exploiting LLMs for Task-Oriented Robot-Human Handovers
- Expanding Computation Spaces of LLMs at Inference Time
- Reasoning or Retrieval? A Study of Answer Attribution on Large Reasoning Models
- ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
- VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
- Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE
- A Small Math Model: Recasting Strategy Choice Theory in an LLM-Inspired Architecture
- AssertFix: Empowering Automated Assertion Fix via Large Language Models
- HiPO: Hybrid Policy Optimization for Dynamic Reasoning in LLMs
- Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm
- EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
- Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
- Navigating the Labyrinth: Path-Sensitive Unit Test Generation with Large Language Models
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
- Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
- Fast Thinking for Large Language Models
- Reasoning Scaffolding: Distilling the Flow of Thought from LLMs
- Improving the Efficiency of LLM Agent Systems through Trajectory Reduction
- Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
- Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
- PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
- Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models
- SAC-Opt: Semantic Anchors for Iterative Correction in Optimization Modeling
- RIV: Recursive Introspection Mask Diffusion Vision Language Model
- Privy: Envisioning and Mitigating Privacy Risks for Consumer-facing AI Product Concepts
- An Senegalese Legal Texts Structuration Using LLM-augmented Knowledge Graph
- Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models
- No Loss, No Gain: Gated Refinement and Adaptive Compression for Prompt Optimization
- MedCritical: Enhancing Medical Reasoning in Small Language Models via Self-Collaborative Correction
- Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought
- Decoupling Reasoning and Perception: An LLM-LMM Framework for Faithful Visual Reasoning
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- Tracing Uncertainty in Language Model "Reasoning"
- GUI-PRA: Process Reward Agent for GUI Tasks
- Self-Consistency as a Free Lunch: Reducing Hallucinations in Vision-Language Models via Self-Reflection
- p-less Sampling: A Robust Hyperparameter-Free Approach for LLM Decoding
- From Harm to Help: Turning Reasoning In-Context Demos into Assets for Reasoning LMs
- Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Space
- Limit Analysis for Symbolic Multi-step Reasoning Tasks with Information Propagation Rules Based on Transformers
- RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobility
- Peacemaker or Troublemaker: How Sycophancy Shapes Multi-Agent Debate
- WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning
- Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned
- Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
- Planning with Unified Multimodal Models
- Causally-Enhanced Reinforcement Policy Optimization
- Towards Human-interpretable Explanation in Code Clone Detection using LLM-based Post Hoc Explainer
- HEART: Emotionally-driven test-time scaling of Language Models
- Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning
- VideoScore2: Think before You Score in Generative Video Evaluation
- What Do They Fix? LLM-Aided Categorization of Security Patches for Critical Memory Bugs
- When Can AI Models Explain Learning? Validity Criteria for AI as Cognitive Models in Education
- UML-CoT: Structured Reasoning and Planning with Unified Modeling Language for Robotic Room Cleaning
- IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
- Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time
- StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models
- Think Socially via Cognitive Reasoning
- The Emergence of Altruism in Large-Language-Model Agents Society
- Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents
- InfiAgent: Self-Evolving Pyramid Agent Framework for Infinite Scenarios
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- MoveFM-R: Advancing Mobility Foundation Models via Language-driven Semantic Reasoning
- A model of errors in transformers
- MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
- Investigating Faithfulness in Large Audio Language Models
- Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models
- R-Capsule: Compressing High-Level Plans for Efficient Large Language Model Reasoning
- SciTS: Scientific Time Series Understanding and Generation with LLMs
- Reinforcement Learning-Guided Chain-of-Draft for Token-Efficient Code Generation
- A2R: An Asymmetric Two-Stage Reasoning Framework for Parallel Reasoning
- Teaching Transformers to Solve Combinatorial Problems through Efficient Trial & Error
- GSM-Agent: Understanding Agentic Reasoning Using Controllable Environments
- Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
- SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
- PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
- Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models
- Unlocking the Essence of Beauty: Advanced Aesthetic Reasoning with Relative-Absolute Policy Optimization
- Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
- Synthetic Dialogue Generation for Interactive Conversational Elicitation & Recommendation (ICER)
- SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
- Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
- Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling
- Retrieval-of-Thought: Efficient Reasoning via Reusing Thoughts
- UISim: An Interactive Image-Based UI Simulator for Dynamic Mobile Environments
- Thinking in Many Modes: How Composite Reasoning Elevates Large Language Model Performance with Limited Data
- Why Chain of Thought Fails in Clinical Text Understanding
- Creative Adversarial Testing (CAT): A Novel Framework for Evaluating Goal-Oriented Agentic AI Systems
- Can AI Perceive Physical Danger and Intervene?
- Towards Transparent AI: A Survey on Explainable Language Models
- Correct Reasoning Paths Visit Shared Decision Pivots
- Plan2Evolve: LLM Self-Evolution for Improved Planning Capability via Automated Domain Generation
- Dual-Head Reasoning Distillation: Improving Classifier Accuracy with Train-Time-Only Reasoning
- Learning to Reason with Mixture of Tokens
- Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data
- Bounds of Chain-of-Thought Robustness: Reasoning Steps, Embed Norms, and Beyond
- Best-of-∞ -- Asymptotic Performance of Test-Time Compute
- Disagreements in Reasoning: How a Model's Thinking Process Dictates Persuasion in Multi-Agent Systems
- A Formal Comparison Between Chain-of-Thought and Latent Thought
- Predicting LLM Reasoning Performance with Small Proxy Model
- Automatic Red Teaming LLM-based Agents with Model Context Protocol Tools
- GALAX: Graph-Augmented Language Model for Explainable Reinforcement-Guided Subgraph Reasoning in Precision Medicine
- StyleBench: Evaluating thinking styles in Large Language Models
- Prompt-Aware Scheduling for Low-Latency LLM Serving
- Stability of In-Context Learning: A Spectral Coverage Perspective
- SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials
- UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
- A Fano-Style Accuracy Upper Bound for LLM Single-Pass Reasoning in Multi-Hop QA
- Difference-Guided Reasoning: A Temporal-Spatial Framework for Large Language Models
- CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
- On Theoretical Interpretations of Concept-Based In-Context Learning
- RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
- d2: Improving Reasoning in Diffusion Language Models via Trajectory Likelihood Estimation
- LayoutAgent: A Vision-Language Agent Guided Compositional Diffusion for Spatial Layout Planning
- MARS: toward more efficient multi-agent collaboration for LLM reasoning
- ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
- (When) Should We Delegate AI Governance to AIs? Some Lessons from Administrative Law
- Thinking Augmented Pre-training
- Automated Multi-Agent Workflows for RTL Design
- Federation of Agents: A Semantics-Aware Communication Fabric for Large-Scale Agentic AI
- The Knowledge-Behaviour Disconnect in LLM-based Chatbots
- GemDetox at TextDetox CLEF 2025: Enhancing a Massively Multilingual Model for Text Detoxification on Low-resource Languages
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- Future Policy Aware Preference Learning for Mathematical Reasoning
- DAOpt: Modeling and Evaluation of Data-Driven Optimization under Uncertainty with LLMs
- MIXRAG : Mixture-of-Experts Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering
- Thinking While Listening: Simple Test Time Scaling For Audio Classification
- SIM-CoT: Supervised Implicit Chain-of-Thought
- SKYLENAGE Technical Report: Mathematical Reasoning and Contest-Innovation Benchmarks for Multi-Level Math Evaluation
- LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning
- Are We Scaling the Right Thing? A System Perspective on Test-Time Scaling
- Mamba Modulation: On the Length Generalization of Mamba
- Semantic-Aware Fuzzing: An Empirical Framework for LLM-Guided, Reasoning-Driven Input Mutation
- SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- Chaos in reason: How chain-of-thought LLMs can look for an answer
- What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
- DeepResearch Agent System
- Scaling LLM-Driven Multi-Agent Systems: Design Principles and Architectural Scalability Analysis
- LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
- Distilling Answer Set Programming Theories from Large Language Models
- OptGraph: Large Language Models Enhanced Evolutionary Optimization Via Graph Retrieval-Augmented Generation
- Can AI Follow In Einstein's Footsteps?
- Challenges in annotations by humans and LLMs: A case study of evaluative language
- Leveraging Trajectory Graphs for Pre-Execution Error Diagnosis in Agentic LLM Systems
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- AutoMem: Automated Learning of Memory as a Cognitive Skill
- Numbers Beat Words: A Rigorous On-Premise Benchmark for Coupled MIMO Controller Tuning
- SuperThoughts: Reasoning Tokens in Superposition
- LASAR: Latent Adaptive Semantic Aligned Reasoning for Generative Recommendation
- AgentClinic: a multimodal benchmark for tool-using clinical AI agents
- Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation
- The Topological Trouble With Transformers
- How ready are we to use artificial intelligence in our fight against antimicrobial resistance? An ESGAID and EAAS perspective
- A Theory of Appropriateness That Accounts for Norms of Rationality
- Non-Parametric Structural Priors for Geometry Theorem Prediction
- Relational Scene Graphs for Object Grounding of Natural Language Commands
- References Improve LLM Alignment in Non-Verifiable Domains
- Automating Conflict-Aware ACL Configurations with Natural Language Intents
- AMELIA: A Family of Multi-task End-to-end Language Models for Argumentation
- FinReflectKG: Agentic Construction and Evaluation of Financial Knowledge Graphs
- InterFeat: a pipeline for finding interesting scientific features
- DRQA: Dynamic Reasoning Quota Allocation for Controlling Overthinking in Reasoning Large Language Models
- CompLLM: Compression for Long Context Q&A
- Agentic Reinforcement Learning with Implicit Step Rewards
- LLMs as verification oracles for Solidity
- On the Soundness and Consistency of LLM Agents for Executing Test Cases Written in Natural Language
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
- Database Normalization via Dual-LLM Self-Refinement
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- Solving Math Word Problems Using Estimation Verification and Equation Generation
- Experience Scaling: Post-Deployment Evolution For Large Language Models
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- Live-E2T: Real-time Threat Monitoring in Video via Deduplicated Event Reasoning and Chain-of-Thought
- Enhancing Automatic Chord Recognition through LLM Chain-of-Thought Reasoning
- Code Driven Planning with Domain-Adaptive Critic
- Advances in Large Language Models for Medicine
- Introducing LongCat-Flash-Thinking: A Technical Report
- GRPO++: Enhancing Dermatological Reasoning under Low Resource Settings
- Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference
- A closed-loop AI framework for hypothesis-driven and interpretable materials design
- iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning
- Harnessing Multimodal Large Language Models for Personalized Product Search with Query-aware Refinement
- Can LLMs Reason Over Non-Text Modalities in a Training-Free Manner? A Case Study with In-Context Representation Learning
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
- AEAS: Actionable Exploit Assessment System
- EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
- MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents
- Mitigating Strategy-Selection Bias in Reasoning for More Effective Test-Time Scaling
- Evaluating the Safety and Skill Reasoning of Large Reasoning Models Under Compute Constraints
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
- Adaptive Kernel Design for Bayesian Optimization Is a Piece of CAKE with LLMs
- ATLAS: Benchmarking and Adapting LLMs for Global Trade via Harmonized Tariff Code Classification
- A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialogue
- OnePiece: Bringing Context Engineering and Reasoning to Industrial Cascade Ranking System
- Automated Facility Enumeration for Building Compliance Checking using Door Detection and Large Language Models
- Agentic AI for Multi-Stage Physics Experiments at a Large-Scale User Facility Particle Accelerator
- From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
- Adaptive Overclocking: Dynamic Control of Thinking Path Length via Real-Time Reasoning Signals
- K-DeCore: Facilitating Knowledge Transfer in Continual Structured Knowledge Reasoning via Knowledge Decoupling
- Difficulty-Aware Score Generation for Piano Sight-Reading
- A Chain-of-thought Reasoning Breast Ultrasound Dataset Covering All Histopathology Categories
- AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
- Large Language Models as End-to-end Combinatorial Optimization Solvers
- MCTS-EP: Empowering Embodied Planning with Online Preference Optimization
- MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
- ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
- RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering
- Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning
- Revisiting Vulnerability Patch Localization: An Empirical Study and LLM-Based Solution
- Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
- SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
- Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
- How Large Language Models are Designed to Hallucinate
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
- How Good are Foundation Models in Step-by-Step Embodied Reasoning?
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human
- Empathy-R1: A Chain-of-Empathy and Reinforcement Learning Framework for Long-Form Mental Health Support
- Enhancing Retrieval Augmentation via Adversarial Collaboration
- Chain-of-Thought Re-ranking for Image Retrieval Tasks
- Evaluating the Impact of Verbal Multiword Expressions on Machine Translation
- TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
- (P)rior(D)yna(F)low: A Priori Dynamic Workflow Construction via Multi-Agent Collaboration
- Large Language Models in Operations Research: Methods, Applications, and Challenges
- Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
- TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
- From Pixels to Urban Policy-Intelligence: Recovering Legacy Effects of Redlining with a Multimodal LLM
- MARIC: Multi-Agent Reasoning for Image Classification
- FlowRL: Matching Reward Distributions for LLM Reasoning
- Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
- From Mimicry to True Intelligence (TI) -- A New Paradigm for Artificial General Intelligence
- PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning
- Combining Evidence and Reasoning for Biomedical Fact-Checking
- THOR: Tool-Integrated Hierarchical Optimization via RL for Mathematical Reasoning
- Thinking in a Crowd: How Auxiliary Information Shapes LLM Reasoning
- DSPC: Dual-Stage Progressive Compression Framework for Efficient Long-Context Reasoning
- DREAM: Domain-aware Reasoning for Efficient Autonomous Underwater Monitoring
- Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
- Early Stopping Chain-of-thoughts in Large Language Models
- VerilogMonkey: Exploring Parallel Scaling for Automated Verilog Code Generation with LLMs
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
- Agentic JWT: A Secure Delegation Protocol for Autonomous AI Agents
- Programmable Cognitive Bias in Social Agents
- AI Agents with Human-Like Collaborative Tools: Adaptive Strategies for Enhanced Problem-Solving
- MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- LATTS: Locally Adaptive Test-Time Scaling
- The Few-shot Dilemma: Over-prompting Large Language Models
- Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs
- Automating Code Generation for Semiconductor Equipment Control from Developer Utterances with LLMs
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- Beyond Data Privacy: New Privacy Risks for Large Language Models
- Leveraging Large Language Models to Effectively Generate Visual Data for Canine Musculoskeletal Diagnoses
- Multi-Robot Task Planning for Multi-Object Retrieval Tasks with Distributed On-Site Knowledge via Large Language Models
- Discovering New Theorems via LLMs with In-Context Proof Learning in Lean
- Analogy-Driven Financial Chain-of-Thought (AD-FCoT): A Prompting Approach for Financial Sentiment Analysis
- Reasoning with Preference Constraints: A Benchmark for Language Models in Many-to-One Matching Markets
- Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
- H2R: Hierarchical Hindsight Reflection for Multi-Task LLM Agents
- Multi-Agent Visual-Language Reasoning for Comprehensive Highway Scene Understanding
- Root Cause Analysis of Radiation Oncology Incidents Using Large Language Models
- Large Language Model-Based Automatic Formulation for Stochastic Optimization Models
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- Text Adaptation to Plain Language and Easy Read via Automatic Post-Editing Cycles
- Prompt Commons: Collective Prompting as Governance for Urban AI
- Evaluating Large Language Models for Functional and Maintainable Code in Industrial Settings: A Case Study at ASML
- HARP: Hallucination Detection via Reasoning Subspace Projection
- Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
- Large language model-empowered next-generation computer-aided engineering
- The Prompt Engineering Report Distilled: Quick Start Guide for Life Sciences
- Free-MAD: Consensus-Free Multi-Agent Debate
- Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
- PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
- The System Description of CPS Team for Track on Driving with Language of CVPR 2024 Autonomous Grand Challenge
- AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
- AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
- Improving Table Understanding with LLMs and Entity-Oriented Search
- ReFineG: Synergizing Small Supervised Models and LLMs for Low-Resource Grounded Multimodal NER
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
- Reasoning Under Uncertainty: Exploring Probabilistic Reasoning Capabilities of LLMs
- CrunchLLM: Multitask LLMs for Structured Business Reasoning and Outcome Prediction
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- RefactorCoderQA: Benchmarking LLMs for Multi-Domain Coding Question Solutions in Cloud and Edge Deployment
- Unveiling the Latent Directions of Reflection in Large Language Models
- WebSight: A Vision-First Architecture for Robust Web Agents
- ReFactX: Scalable Reasoning with Reliable Facts via Constrained Generation
- Who Decides How Knowing Becomes Doing? Redistributing Authority in Human-AI Music Co-Creation
- Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Development of Automated Software Design Document Review Methods Using Large Language Models
- Smart Trial: Evaluating the Use of Large Language Models for Recruiting Clinical Trial Participants via Social Media
- InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- Retrieval-Augmented Generation for Reliable Interpretation of Radio Regulations
- MetaLLMix : An XAI Aided LLM-Meta-learning Based Approach for Hyper-parameters Optimization
- Modelling Analogies and Analogical Reasoning: Connecting Cognitive Science Theory and NLP Research
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- Harnessing Uncertainty: Entropy-Modulated Policy Gradients for Long-Horizon LLM Agents
- Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining
- How well can LLMs provide planning feedback in grounded environments?
- SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds
- LLMs as Agentic Cooperative Players in Multiplayer UNO
- PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability
- Evaluating LLMs Without Oracle Feedback: Agentic Annotation Evaluation Through Unsupervised Consistency Signals
- The meaning of prompts and the prompts of meaning: Semiotic reflections and modelling
- An Iterative LLM Framework for SIBT utilizing RAG-based Adaptive Weight Optimization
- HyperMOOC: Augmenting MOOC Videos with Concept-based Embedded Visualizations
- FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
- Gala: Global LLM Agents for Text-to-Model Translation
- Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models
- Ubiquitous Intelligence Via Wireless Network-Driven LLMs Evolution
- Transparency of medical artificial intelligence systems
- Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- BREATH: A Bio-Radar Embodied Agent for Tonal and Human-Aware Diffusion Music Generation
- Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)
- AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
- Graph-Fused Vision-Language-Action for Policy Reasoning in Multi-Arm Robotic Manipulation
- Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
- Certainty-Guided Reasoning in Large Language Models: A Dynamic Thinking Budget Approach
- From Detection to Mitigation: Addressing Gender Bias in Chinese Texts via Efficient Tuning and Voting-Based Rebalancing
- Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- Reconstruction Alignment Improves Unified Multimodal Models
- HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
- Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning
- Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
- RAFFLES: Reasoning-based Attribution of Faults for LLM Systems
- Small Vectors, Big Effects: A Mechanistic Study of RL-Induced Reasoning via Steering Vectors
- Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
- Teaching AI Stepwise Diagnostic Reasoning with Report-Guided Chain-of-Thought Learning
- A Fragile Number Sense: Probing the Elemental Limits of Numerical Reasoning in LLMs
- Can AI Make Energy Retrofit Decisions? An Evaluation of Large Language Models
- The Thinking Therapist: Training Large Language Models to Deliver Acceptance and Commitment Therapy using Supervised Fine-Tuning and Odds Ratio Policy Optimization
- The Majority is not always right: RL training for solution aggregation
- PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
- Text-Trained LLMs Can Zero-Shot Extrapolate PDE Dynamics, Revealing a Three-Stage In-Context Learning Mechanism
- MedBench-IT: A Comprehensive Benchmark for Evaluating Large Language Models on Italian Medical Entrance Examinations
- Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
- Rethinking LLM Parametric Knowledge as Post-retrieval Confidence for Dynamic Retrieval and Reranking
- Language Native Lightly Structured Databases for Large Language Model Driven Composite Materials Research
- Benchmarking Information Retrieval Models on Complex Retrieval Tasks
- From Digital Distrust to Codified Honesty: Experimental Evidence on Generative AI in Credence Goods Markets
- Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial
- GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation
- Stack Overflow Is Not Dead Yet: Crowd Answers Still Matter
- PolicyEvolve: Evolving Programmatic Policies by LLMs for multi-player games via Population-Based Training
- Reverse-Engineered Reasoning for Open-Ended Generation
- Reasoning Language Model for Personalized Lung Cancer Screening
- From Long to Short: LLMs Excel at Trimming Own Reasoning Chains
- ProfilingAgent: Profiling-Guided Agentic Reasoning for Adaptive Model Optimization
- Cross-Question Method Reuse in Large Language Models: From Word-Level Prediction to Rational Logical-Layer Reasoning
- AI-Assisted Modeling: DSL-Driven AI Interactions
- BEDTime: A Unified Benchmark for Automatically Describing Time Series
- Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
- What-If Analysis of Large Language Models: Explore the Game World Using Proactive Thinking
- Less is More Tokens: Efficient Math Reasoning via Difficulty-Aware Chain-of-Thought Distillation
- Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
- Towards Open World Detection: A Survey
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- CoVeR: Conformal Calibration for Versatile and Reliable Autoregressive Next-Token Prediction
- Hunyuan-MT Technical Report
- TAGAL: Tabular Data Generation using Agentic LLM Methods
- Characterizing Fitness Landscape Structures in Prompt Engineering
- Intermediate Languages Matter: Formal Languages and LLMs affect Neurosymbolic Reasoning
- LLM-as-classifier: Semi-Supervised, Iterative Framework for Hierarchical Text Classification using Large Language Models
- Finance-Grounded Optimization For Algorithmic Trading
- CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning
- On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
- Long-Horizon Visual Imitation Learning via Plan and Code Reflection
- FaMA: LLM-Empowered Agentic Assistant for Consumer-to-Consumer Marketplace
- Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Expedition & Expansion: Leveraging Semantic Representations for Goal-Directed Exploration in Continuous Cellular Automata
- PromptEnhancer: A Simple Approach to Enhance Text-to-Image Models via Chain-of-Thought Prompt Rewriting
- Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
- Explainable Knowledge Graph Retrieval-Augmented Generation (KG-RAG) with KG-SMILE
- Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents
- sam-llm: interpretable lane change trajectoryprediction via parametric finetuning
- Generative Auto-Bidding in Large-Scale Competitive Auctions via Diffusion Completer-Aligner
- PromptCOS: Towards Content-only System Prompt Copyright Auditing for LLMs
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Are We SOLID Yet? An Empirical Study on Prompting LLMs to Detect Design Principle Violations
- Plan More, Debug Less: Applying Metacognitive Theory to AI-Assisted Programming Education
- OPRA-Vis: Visual Analytics System to Assist Organization-Public Relationship Assessment with Large Language Models
- Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
- Towards Reasoning for PDE Foundation Models: A Reward-Model-Driven Inference-Time-Scaling Algorithm
- OwkinZero: Accelerating Biological Discovery with AI
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
- Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation
- Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks
- Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety
- Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time
- AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
- A QoE-Driven Personalized Incentive Mechanism Design for AIGC Services in Resource-Constrained Edge Networks
- Better by Comparison: Retrieval-Augmented Contrastive Reasoning for Automatic Prompt Optimization
- Generative KI für TA
- Language-Guided Long Horizon Manipulation with LLM-based Planning and Visual Perception
- Batch Query Processing and Optimization for Agentic Workflows
- LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents
- Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving
- DynaGuard: A Dynamic Guardian Model With User-Defined Policies
- Extracting OPQRST in Electronic Health Records using Large Language Models with Reasoning
- Graph RAG as Human Choice Model: Building a Data-Driven Mobility Agent with Preference Chain
- Baichuan-M2: Scaling Medical Capability with Large Verifier System
- Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models
- Dynamic Speculative Agent Planning
- Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
- Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
- Relative Trajectory Balance is equivalent to Trust-PCL
- Strata-Sword: A Hierarchical Safety Evaluation towards LLMs based on Reasoning Complexity of Jailbreak Instructions
- Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
- Rethinking the Chain-of-Thought: The Roles of In-Context Learning and Pre-trained Priors
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- Generative Goal Modeling
- Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- An Automated Attack Investigation Approach Leveraging Threat-Knowledge-Augmented Large Language Models
- TableZoomer: A Collaborative Agent Framework for Large-scale Table Question Answering
- Exploring and Mitigating Fawning Hallucinations in Large Language Models
- Aligning Reasoning LLMs for Materials Discovery with Physics-aware Rejection Sampling
- CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs
- Decomposing and Revising What Language Models Generate
- On Verifiable Legal Reasoning: A Multi-Agent Framework with Formalized Knowledge Representations
- Designing LMS and Instructional Strategies for Integrating Generative-Conversational AI
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- VULSOLVER: Vulnerability Detection via LLM-Driven Constraint Solving
- RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
- One VLM, Two Roles: Stage-Wise Routing and Specialty-Level Deployment for Clinical Workflows
- When Thinking Backfires: Mechanistic Insights Into Reasoning-Induced Misalignment
- Memory Limitations of Prompt Tuning in Transformers
- Open Data Synthesis For Deep Research
- Access Paths for Efficient Ordering with Large Language Models
- ConceptBot: Enhancing Robot's Autonomy through Task Decomposition with Large Language Models and Knowledge Graph
- Illuminating Patterns of Divergence: DataDios SmartDiff for Large-Scale Data Difference Analysis
- Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- DriveQA: Passing the Driving Knowledge Test
- PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
- Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight
- How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images
- Failure Prediction Is a Better Performance Proxy for Early-Exit Networks Than Calibration
- AI Compute Architecture and Evolution Trends
- The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
- Evaluation of Large Language Models for Anomaly Detection in Autonomous Vehicles
- From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics
- Zero‐ and few‐shot prompting of generative large language models provides weak assessment of risk of bias in clinical trials
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- Improving Aviation Safety Analysis: Automated HFACS Classification Using Reinforcement Learning with Group Relative Policy Optimization
- An Agile Method for Implementing Retrieval Augmented Generation Tools in Industrial SMEs
- ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
- WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
- Re4: Scientific Computing Agent with Rewriting, Resolution, Review and Revision
- Language-Enhanced Mobile Manipulation for Efficient Object Search in Indoor Environments
- Bridging Minds and Machines: Toward an Integration of AI and Cognitive Science
- AI-SearchPlanner: Modular Agentic Search via Pareto-Optimal Multi-Objective Reinforcement Learning
- MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning
- Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
- Provable Benefits of In-Tool Learning for Large Language Models
- Adaptive Root Cause Localization for Microservice Systems with Multi-Agent Recursion-of-Thought
- Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
- Automated Quality Assessment for LLM-Based Complex Qualitative Coding: A Confidence-Diversity Framework
- Symphony: A Decentralized Multi-Agent Framework for Scalable Collective Intelligence
- SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control
- CataractSurg-80K: Knowledge-Driven Benchmarking for Structured Reasoning in Ophthalmic Surgery Planning
- Evaluating Language Model Reasoning about Confidential Information
- PersoNo: Personalised Notification Urgency Classifier in Mixed Reality
- CompLex: Music Theory Lexicon Constructed by Autonomous Agents for Automatic Music Generation
- Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
- AI reasoning effort predicts human decision time in content moderation
- Think in Blocks: Adaptive Reasoning from Direct Response to Deep Reasoning
- Applications of GPT in Political Science Research: Extracting Information from Unstructured Text
- GENIE-ASI: Generative Instruction and Executable Code for Analog Subcircuit Identification
- LongReasonArena: A Long Reasoning Benchmark for Large Language Models
- Autoregressive Universal Video Segmentation Model
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- StepWiser: Stepwise Generative Judges for Wiser Reasoning
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning
- Retrieval-Augmented Review Generation for Poisoning Recommender Systems
- Reasoning LLMs in the Medical Domain: A Literature Survey
- ArgRAG: Explainable Retrieval Augmented Generation using Quantitative Bipolar Argumentation
- ReflectivePrompt: Reflective evolution in autoprompting algorithms
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision
- Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
- Language and Experience: A Computational Model of Social Learning in Complex Tasks
- VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft
- Uncovering Intervention Opportunities for Suicide Prevention with Language Model Assistants
- Language Models For Generalised PDDL Planning: Synthesising Sound and Programmatic Policies
- VERIRL: Boosting the LLM-based Verilog Code Generation via Reinforcement Learning
- SafeBimanual: Diffusion-based Trajectory Optimization for Safe Bimanual Manipulation
- Type-Compliant Adaptation Cascades: Adapting Programmatic LM Workflows to Data
- InReAcTable: LLM-Powered Interactive Visual Data Story Construction from Tabular Data
- Test-Time Scaling Strategies for Generative Retrieval in Multimodal Conversational Recommendations
- Scaling Group Inference for Diverse and High-Quality Generation
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Deep Think with Confidence
- R-ConstraintBench: Evaluating LLMs on NP-Complete Scheduling
- Fine-tuning and prompt engineering for large language models-based code review automation
- Long Chain-of-Thought Reasoning Across Languages
- From Sound to Sight: Towards AI-authored Music Videos
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
- In2x at WMT25 Translation Task
- MedCoT-RAG: Causal Chain-of-Thought RAG for Medical Question Answering
- Static Analysis as a Feedback Loop: Enhancing LLM-Generated Code Beyond Correctness
- Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
- PepThink-R1: LLM for Interpretable Cyclic Peptide Optimization with CoT SFT and Reinforcement Learning
- The Prompting Brain: Neurocognitive Markers of Expertise in Guiding Large Language Models
- LLMs and Agentic AI in Insurance Decision-Making: Opportunities and Challenges For Africa
- Lexical Hints of Accuracy in LLM Reasoning Chains
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- Jointly Extracting Interventions, Outcomes, and Findings from RCT Reports with LLMs
- LLM-Powered Virtual Patient Agents for Interactive Clinical Skills Training with Automated Feedback
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- CausalPlan: Empowering Efficient LLM Multi-Agent Collaboration Through Causality-Driven Planning
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- Input-Time Scaling
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
- Structured Prompting and Multi-Agent Knowledge Distillation for Traffic Video Interpretation and Risk Inference
- Machine learning in fluid dynamics: A critical assessment
- COCO: Cognitive Operating System with Continuous Oversight for Multi-Agent Workflow Reliability
- Driving Style Recognition Like an Expert Using Semantic Privileged Information from Large Language Models
- Datarus-R1: An Adaptive Multi-Step Reasoning LLM for Automated Data Analysis
- AI Agents for Photonic Integrated Circuit Design Automation
- Reinforced Context Order Recovery for Adaptive Reasoning and Planning
- CardAIc-Agents: A Multimodal Framework with Hierarchical Adaptation for Cardiac Care Support
- G2RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance
- Involuntary Jailbreak: On Self-Prompting Attacks
- Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
- CRED-SQL: Enhancing Real-world Large Scale Database Text-to-SQL Parsing through Cluster Retrieval and Execution Description
- Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning
- Enhancing Cryptocurrency Sentiment Analysis with Multimodal Features
- RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
- DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning
- Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
- Leveraging Large Language Models for Predictive Analysis of Human Misery
- Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
- OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
- Can Large Models Teach Student Models to Solve Mathematical Problems Like Human Beings? A Reasoning Distillation Method via Multi-LoRA Interaction
- Adaptive Reinforcement for Open-ended Medical Reasoning via Semantic-Guided Reward Collapse Mitigation
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Uncovering Emergent Physics Representations Learned In-Context by Large Language Models
- Hierarchical knowledge guided fault intensity diagnosis of complex industrial systems
- EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos
- Synchronization Dynamics of Heterogeneous, Collaborative Multi-Agent AI Systems
- Fast, Slow, and Tool-augmented Thinking for LLMs: A Review
- Disentangling the Drivers of LLM Social Conformity: An Uncertainty-Moderated Dual-Process Mechanism
- Improving Pre-Trained Vision-Language-Action Policies with Model-Based Search
- Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow
- Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
- GraphCogent: Mitigating LLMs' Working Memory Constraints via Multi-Agent Collaboration in Complex Graph Understanding
- Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
- Exploring Efficiency Frontiers of Thinking Budget in Medical Reasoning: Scaling Laws between Computational Resources and Reasoning Quality
- Mitigating Jailbreaks with Intent-Aware LLMs
- Benchmarking LLM-based Agents for Single-cell Omics Analysis
- QuarkMed Medical Foundation Model Technical Report
- LARC: Towards Human-level Constrained Retrosynthesis Planning through an Agentic Framework
- A Comprehensive Review of AI Agents: Transforming Possibilities in Technology and Beyond
- ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
- A Multi-Task Evaluation of LLMs' Processing of Academic Text Input
- CryptoScope: Utilizing Large Language Models for Automated Cryptographic Logic Vulnerability Detection
- Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models
- SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling
- Defects4Log: Benchmarking LLMs for Logging Code Defect Detection and Reasoning
- Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework
- HumorPlanSearch: Structured Planning and HuCoT for Contextual AI Humor
- Retrieval-augmented reasoning with lean language models
- Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
- MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
- Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
- WIP: Leveraging LLMs for Enforcing Design Principles in Student Code: Analysis of Prompting Strategies and RAG
- Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
- SSRL: Self-Search Reinforcement Learning
- UI-Venus Technical Report: Building High-performance UI Agents with RFT
- Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rules
- FROGENT: An End-to-End Full-process Drug Design Agent
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- X-Node: Self-Explanation is All We Need
- We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
- Computational Economics in Large Language Models: Exploring Model Behavior and Incentive Design under Resource Constraints
- Evaluating LLMs on Chinese Idiom Translation
- A Semantic-Aware Framework for Safe and Intent-Integrative Assistance in Upper-Limb Exoskeletons
- MCP-Enabled LLM for Meta-optics Inverse Design: Leveraging Differentiable Solver without LLM Expertise
- Ask ChatGPT: Caveats and Mitigations for Individual Users of AI Chatbots
- MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
- LingVarBench: Benchmarking LLM for Automated Named Entity Recognition in Structured Synthetic Spoken Transcriptions
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
- Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development
- Apriel-Nemotron-15B-Thinker
- Exploring the Potential of Large Language Models in Fine-Grained Review Comment Classification
- Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study
- ReqInOne: A Large Language Model-Based Agent for Software Requirements Specification Generation
- Towards Self-cognitive Exploration: Metacognitive Knowledge Graph Retrieval Augmented Generation
- A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report Generation
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- Cross-lingual Aspect-Based Sentiment Analysis: A Survey on Tasks, Approaches, and Challenges
- Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation
- LLM Empowered Prototype Learning for Zero and Few-Shot Tasks on Tabular Data
- Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages
- FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
- How Does a Virtual Agent Decide Where to Look? -- Symbolic Cognitive Reasoning for Embodied Head Rotation
- The Roots of International Perceptions: Simulating US Attitude Changes Towards China with LLM Agents
- A Survey of Optimization Modeling Meets LLMs: Progress and Future Directions
- Can We Trust AI to Govern AI? Benchmarking LLM Performance on Privacy and AI Governance Exams
- Exploring Large Language Model Agents for Piloting Social Experiments
- Agentic Graph Neural Networks for Wireless Communications and Networking Towards Edge General Intelligence: A Survey
- DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
- Leveraging Large Language Models for Rare Disease Named Entity Recognition
- InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling
- Rational Inverse Reasoning: Few-Shot Imitation by Inferring Intent through Planning
- DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
- Train Long, Think Short: Curriculum Learning for Efficient Reasoning
- Time Is a Feature: Exploiting Temporal Dynamics in Diffusion Language Models
- MuaLLM: A Multimodal Large Language Model Agent for Circuit Design Assistance with Hybrid Contextual Retrieval-Augmented Generation
- WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
- \(X\)-evolve: Solution space evolution powered by large language models
- Evaluating Large Language Models as Expert Annotators
- Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
- CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts
- Retrieval-Augmented Generation in Industry: An Interview Study on Use Cases, Requirements, Challenges, and Evaluation
- Large Language Models for Subjective Language Understanding: A Survey
- TeamMedAgents: Pareto-Efficient Multi-Agent Medical Reasoning Through Teamwork Theory
- Keyword-Centric Prompting for One-Shot Event Detection with Self-Generated Rationale Enhancements
- FEAT: A Multi-Agent Forensic AI System with Domain-Adapted Large Language Model for Automated Cause-of-Death Analysis
- Grid2Guide: A* Enabled Small Language Model for Indoor Navigation
- In-situ Value-aligned Human-Robot Interactions with Physical Constraints
- MolmoAct: Action Reasoning Models that can Reason in Space
- Vision-Based Localization and LLM-based Navigation for Indoor Environments
- MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
- Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models
- CP-Agent: Agentic Constraint Programming
- A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
- LLM-based Agents for Automated Confounder Discovery and Subgroup Analysis in Causal Inference
- Schema Lineage Extraction at Scale: Multilingual Pipelines, Composite Evaluation, and Language-Model Benchmarks
- How Effectively Can Large Language Models Connect SNP Variants and ECG Phenotypes for Cardiovascular Risk Prediction?
- Arce: Augmented Roberta with Contextualized Elucidations for Ner in Automated Rule Checking
- Positional Biases Shift as Inputs Approach Context Window Limits
- Uncertainty-Aware Semantic Decoding for LLM-Based Sequential Recommendation
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Story Ribbons: Reimagining Storyline Visualizations with Large Language Models
- White-Box Reasoning: Synergizing LLM Strategy and gm/Id Data for Automated Analog Circuit Design
- Integrating Rules and Semantics for LLM-Based C-to-Rust Translation
- Towards High-Order Mean Flow Generative Models: Feasibility, Expressivity, and Provably Efficient Criteria
- VL-MedGuide: A Visual-Linguistic Large Model for Intelligent and Explainable Skin Disease Auxiliary Diagnosis
- End-to-End Text-to-SQL with Dataset Selection: Leveraging LLMs for Adaptive Query Generation
- From Explainable to Explanatory Artificial Intelligence: Toward a New Paradigm for Human-Centered Explanations through Generative AI
- PanelTR: Zero-Shot Table Reasoning Framework Through Multi-Agent Scientific Discussion
- Devstral: Fine-tuning Language Models for Coding Agent Applications
- Optimizing Prompt Sequences using Monte Carlo Tree Search for LLM-Based Optimization
- First Ask Then Answer: A Framework Design for AI Dialogue Based on Supplementary Questioning with Large Language Models
- Do Biased Models Have Biased Thoughts?
- Learning by Teaching: Engaging Students as Instructors of Large Language Models in Computer Science Education
- Large Language Models for Oral History Understanding with Text Classification and Sentiment Analysis
- Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- AI vs. Human Moderators: A Comparative Evaluation of Multimodal LLMs in Content Moderation for Brand Safety
- Leveraging LLMs for Privacy-Aware Predictions in Participatory Budgeting
- Can Language Models Critique Themselves? Investigating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025
- Information-Theoretic Graph Fusion with Vision-Language-Action Model for Policy Reasoning and Dual Robotic Control
- Textual Inversion for Efficient Adaptation of Open-Vocabulary Object Detectors Without Forgetting
- A Novel Architecture for Symbolic Reasoning with Decision Trees and LLM Agents
- An Explainable Natural Language Framework for Identifying and Notifying Target Audiences In Enterprise Communication
- ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
- Incident Response Planning Using a Lightweight Large Language Model with Reduced Hallucination
- Evaluation of Finetuned LLMs in AMR Parsing
- A Multi-Stage Large Language Model Framework for Extracting Suicide-Related Social Determinants of Health
- Charts-of-Thought: Enhancing LLM Visualization Literacy Through Structured Data Extraction
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
- Chain of Questions: Guiding Multimodal Curiosity in Language Models
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
- Deliberative Reasoning Network: An Uncertainty-Driven Paradigm for Belief-Tracked Inference with Pretrained Language Models
- Method-Based Reasoning for Large Language Models: Extraction, Reuse, and Continuous Improvement
- Efficient Strategy for Improving Large Language Model (LLM) Capabilities
- KG-Augmented Executable CoT for Mathematical Coding
- AttriLens-Mol: Attribute Guided Reinforcement Learning for Molecular Property Prediction with Large Language Models
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
- ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
- GeoSR: Cognitive-Agentic Framework for Probing Geospatial Knowledge Boundaries via Iterative Self-Refinement
- AttnTrace: Attention-based Context Traceback for Long-Context LLMs
- Self-Questioning Language Models
- FairLangProc: A Python package for fairness in NLP
- MultiRAG: A Knowledge-guided Framework for Mitigating Hallucination in Multi-source Retrieval Augmented Generation
- CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation
- Unravelling the Probabilistic Forest: Arbitrage in Prediction Markets
- Agoran: An Agentic Open Marketplace for 6G RAN Automation
- A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning
- Compressing Chain-of-Thought in LLMs via Step Entropy
- From Legacy to Standard: LLM-Assisted Transformation of Cybersecurity Playbooks into CACAO Format
- NLP Methods May Actually Be Better Than Professors at Estimating Question Difficulty
- CoTox: Chain-of-Thought-Based Molecular Toxicity Reasoning and Prediction
- When AI Evaluates Its Own Work: Validating Learner-Initiated, AI-Generated Physics Practice Problems
- On the Evaluation of Large Language Models in Multilingual Vulnerability Repair
- Künstliche Intelligenz in den Naturwissenschaftsdidaktiken – gekommen, um zu bleiben: Potenziale, Desiderata, Herausforderungen
- CTTS: Collective Test-Time Scaling
- Seemingly Simple Planning Problems are Computationally Challenging: The Countdown Game
- Cognitive Loop via In-Situ Optimization: Self-Adaptive Reasoning for Science
- CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
- AnalogCoder-Pro: Unifying Analog Circuit Generation and Optimization via Multi-modal LLMs
- XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models
- A Confidence-Diversity Framework for Calibrating AI Judgement in Accessible Qualitative Coding Tasks
- MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
- LMAR: Language Model Augmented Retriever for Domain-specific Knowledge Indexing
- SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents
- ReflecSched: Solving Dynamic Flexible Job-Shop Scheduling via LLM-Powered Hierarchical Reflection
- DeepVIS: Bridging Natural Language and Data Visualization Through Step-wise Reasoning
- Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention
- Towards Zero-Shot Terrain Traversability Estimation: Challenges and Opportunities
- Tuning LLM-based Code Optimization via Meta-Prompting: An Industrial Perspective
- MedSynth: Realistic, Synthetic Medical Dialogue-Note Pairs
- Prompting Large Language Models with Partial Knowledge for Answering Questions with Unseen Entities
- Bridging LLMs and Symbolic Reasoning in Educational QA Systems: Insights from the XAI Challenge at IJCNN 2025
- Transformers in Pseudo-Random Number Generation: A Dual Perspective on Theory and Practice
- Effective Damage Data Generation by Fusing Imagery with Human Knowledge Using Vision-Language Models
- ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models
- Large language models accurately identify immunosuppression in intensive care unit patients
- Exploring Direct Instruction and Summary-Mediated Prompting in LLM-Assisted Code Modification
- The Promise of RL for Autoregressive Image Editing
- Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language Models
- DAMR: Efficient and Adaptive Context-Aware Knowledge Graph Question Answering with LLM-Guided MCTS
- MCeT: Behavioral Model Correctness Evaluation using Large Language Models
- Thinking Machines: Mathematical Reasoning in the Age of LLMs
- Blueprint First, Model Second: A Framework for Deterministic LLM Workflow
- WMAS: A Multi-Agent System Towards Intelligent and Customized Wireless Networks
- Multimodal Referring Segmentation: A Survey
- Accurate and Consistent Graph Model Generation from Text with Large Language Models
- Diagnostic Accuracy of Open-Source Vision-Language Models on Diverse Medical Imaging Tasks
- Lucy: edgerunning agentic web search on mobile with machine generated task vectors
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
- Exploring the Role of LLMs Like ChatGPT in Pharmacy Education for Supporting Students' Therapeutic Decision-making
- EMA Without the Lag: Bias-Corrected Iterate Averaging Schemes
- SimuRA: A World-Model-Driven Simulative Reasoning Architecture for General Goal-Oriented Agents
- BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
- LENS: Learning Ensemble Confidence from Neural States for Multi-LLM Answer Integration
- How Far Are AI Scientists from Changing the World?
- Rule2Text: Natural Language Explanation of Logical Rules in Knowledge Graphs
- Can LLM-Reasoning Models Replace Classical Planning? A Benchmark Study
- MemoCue: Empowering LLM-Based Agents for Human Memory Recall via Strategy-Guided Querying
- MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models
- Causal Reasoning in Pieces: Modular In-Context Learning for Causal Discovery
- Distributed AI Agents for Cognitive Underwater Robot Autonomy
- DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
- TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning
- The Multi-Agent Fault Localization System Based on Monte Carlo Tree Search Approach
- GPT-4.1 Sets the Standard in Automated Experiment Design Using Novel Python Libraries
- Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
- MetaAgent: Automatically Constructing Multi-Agent Systems Based on Finite State Machines
- RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment
- Visual Language Models as Zero-Shot Deepfake Detectors
- Towards Interpretable Renal Health Decline Forecasting via Multi-LMM Collaborative Reasoning Framework
- Exploring In-Context Learning for Frame-Semantic Parsing
- LLMs Between the Nodes: Community Discovery Beyond Vectors
- Multimodal Video Emotion Recognition with Reliable Reasoning Priors
Discussions
- Chain of Thought (CoT) reasoning in multi-agentic #AI helps different AI agents solve problems step-by-step by sharing their thoughts and progress, similar to humans discussing and breaking down compl [bsky, 5 points, 0 comments]
- Chain of Thought Prompting Elicits Reasoning in Large Language Models [hn, 2 points, 1 comments]
- New Apple study challenges whether AI models truly “reason” through problems [lemmy, 1 points, 0 comments]
- Google paper's arxiv.org/pdf/2201.11903. shows that simply asking large language models to "think step by step" dramatically improves their reasoning abilities. Researchers found this technique boosts [bsky, 1 points, 0 comments]
- Chain of Thought Prompting Elicits Reasoning in Large Language Models [hn, 1 points, 0 comments]
- 🎧 EP034: Chain of Thought Prompting Unlocks Reasoning 📄 Chain-of-Thought 🔗 https://arxiv.org/abs/2201.11903 🟢 https://podcasters.spotify.com/pod/show/yun-wu/episodes/EP034-Chain-of-Thought-Prompti [bsky, 0 points, 0 comments]
- The chain of thought prompting technique conclusively improves results and that's just one successful example of prompt engineering. arxiv.org/abs/2201.11903 [bsky, 0 points, 0 comments]
- Chain-of-Thought paper
arxiv.org/abs/2201.11903 [bsky, 0 points, 1 comments]
- 2/works. Am not an AI expert.
There are three broad categories of how DS works. 1) Chain of thought (CT) 2) Reinforcement learning 3) Model distillation.
1) Here's the arxiv paper on CT:
arxiv.org/ [bsky, 0 points, 1 comments]
- arxiv.org/abs/2201.11903 arxiv.org/abs/2304.03439 www.sciencedirect.com/science/arti... www.cnet.com/tech/service... [bsky, 0 points, 0 comments]
- It can. It was capable of this pre-2024, but Chain of Thought significantly increased its abilities to conduct multi-step problem-solving. This 2022 paper was influential: arxiv.org/abs/2201.11903 But [bsky, 0 points, 1 comments]
- arxiv.org/abs/2201.11903 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2201.11903 [bsky, 0 points, 1 comments]
- This is just Chain of Thought prompting! Plus they probably trained it on examples of CoT. arxiv.org/abs/2201.11903 [bsky, 0 points, 0 comments]
- (´-`).。oO( CoTのアイディアも最初はGoogleが出した arxiv.org/abs/2201.11903 のに,アテンションモデル/Transformerと同様に実装ではOpenAIに抜かされちゃうの,既視感しかないというか… ) [bsky, 0 points, 0 comments]
- Interesting idea! I'd separate the concepts of chain-of-thought (and generally, inference-time scaling) from the UX needs.
The chain of thought is literally required to produce the result (see the or [bsky, 0 points, 1 comments]
- arxiv.org/pdf/2201.11903 [bsky, 0 points, 0 comments]
- [2201.11903] Chain-of-Thought Prompting Elicits Reasoning in Large Language Models arxiv.org/abs/2201.11903 [bsky, 0 points, 0 comments]
Related