Training language models to follow instructions with human feedback
2022/03/04 by Long Ouyang, Ouyang, Long, Jeff Wu +38 · 9 voices · 2460 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Explainable Artificial Intelligence (XAI)
paper · pdf · doi:10.48550/arxiv.2203.02155
Abstract
Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users. In this paper, we show an avenue for aligning language models with user intent on a wide range of tasks by fine-tuning with human feedback. Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback. We call the resulting models InstructGPT. In human evaluations on our prompt distribution, outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. Moreover, InstructGPT models show improvements in truthfulness and reductions in toxic output generation while having minimal performance regressions on public NLP datasets. Even though InstructGPT still makes simple mistakes, our results show that fine-tuning with human feedback is a promising direction for aligning language models with human intent.
Cited by
- ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents
- Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
- When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
- Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
- Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs
- Unified Static-Dynamic Pruning for Efficient LLM Inference
- Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
- Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models
- Three-Body Alignment: Aligning Chess Agent with Human Reasoning through Reranked Rationale
- Hint-Guided Diversified Policy Optimization for LLM Reasoning
- ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents
- Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives
- Improving Large Vision-Language Models' Understanding for Flow Field Data
- From Mind to Machine: The Rise of Manus AI as a Fully Autonomous Digital Agent
- Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
- Learning to Reason for Factuality
- Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability
- Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints
- A Unified Moral-Value Dataset for Instruction Tuning
- PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
- Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning
- Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
- Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner
- PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
- SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning
- RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
- AlphaRoute: Large Language Models as Semantic Optimizers for Multi-Objective Routing
- PercepCap: Video Captioner with Structured Spatio-Temporal Perception
- Towards Automated Formal Verification of zkEVMs Using LLM-Guided Constraint Synthesis
- Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
- Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design
- Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents
- Sound Probabilistic Safety Bounds for Large Language Models
- SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement
- Verifiable Self-Evolution for Open-Ended Dialogue Skills via Future-Feedback Prediction
- Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation
- Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
- Post-Training in Time Series Foundation Models: A Unifying Framework
- Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment
- SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking
- REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning
- Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks
- Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency
- How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
- XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
- Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models
- Task Competence Is Not Instruction Following: Evaluating Instruction-Conflicting Behavior in Small Language Models
- When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play
- Semantic-Aware Data-Aided Channel Estimation with Large Language Models for MIMO Systems
- Rushes: A Human Preference Dataset for Pluralistic Alignment
- Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
- Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels
- Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models
- Rater State Bias in RLHF Preference Data: An Audit Framework
- Dr. Zero: Self-Evolving Search Agents without Training Data
- RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
- LP-SFT: Local-Preserving Supervised Fine-Tuning via Multimodal Entropy Structure
- Falsifiable Release Gates for Self-Improving Systems: Standing Invariants at Scale
- In-Context Learning for Wound Classification with Small Multimodal Language Models
- PlotTwist: A Creative Plot Generation Framework with Small Language Models
- Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs
- DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization
- Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction
- Discovery by Dreaming: Cross-Domain Recombination in Artificial Memory
- Meta-Learning Preferences for Multilingual LLM Alignment
- Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
- LLM-Driven Cross-Paradigm Design for Quantum Optimal Control
- ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples
- Matching Ranks Over Probability Yields Truly Deep Safety Alignment
- After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions
- BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
- MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
- A Geometric Perspective on Stabilizing Value Conflict Resolution
- Signed Rectified Flow: Negativity-Controlled Generation
- LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications
- OR Else: A Differentiable Trust Region for Policy Optimization
- Assisting or resisting patriarchy? a critical discourse analysis of chatgpt’s responses on feminism
- ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
- RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts
- Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
- DeLIVeR: Decomposed Learning for Information-grounded Veracity Recognition via Reinforced Knowledge Graph Exploration
- AI Value Alignment for Evolving Social Norms
- Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
- Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
- Rationalizing Boltzmann Rationality: An Axiomatic Characterization of Entropy-Regularized Policies
- Distilled Reinforcement Learning for LLM Post-training
- Trace-Based On-Policy Distillation for Masked Diffusion Language Models
- The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
- The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible
- Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL
- Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis
- LogicIF: Towards Complex Logic Instruction Following
- Model-Driven Discipline for Multi-Agent LLMs: Requirement-to-Verification Generation of Traceable System Models
- RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
- Diversity-Oriented Fine-Tuning for Uncertainty-Based Hallucination Detection
- Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
- Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets
- Scaling Point-in-Time Language Models
- Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
- TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue
- RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation
- Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
- Decoupled Alignment for Robust Plug-and-Play Adaptation
- Nonuniformity Principle in Human-AI Coworking
- One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models
- Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function
- Auditing Inference-Time Defense Evaluation for Multimodal Large Language Models
- When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
- Relational Preference Encoding in Looped Transformer Internal States
- Understanding Agent-Reactive Bugs at the Model-Harness Boundary: An Empirical Study of LLM Agent Issue Reports
- The CRAFT principles for the responsible use of large language models in policymaking
- EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
- Mask-Aware Policy Gradients for Diffusion Language Models
- Step-Level Preference Learning for Generative Agents in Social Simulations
- DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
- Multi-Turn On-Policy Distillation with Prefix Replay
- Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
- CoTu at EXACT 2026: Neuro-Symbolic Reasoning for Transparent Educational QA
- Scaling Evaluation-time Compute with Reasoning Models as Evaluators
- From Stateless to Situated: Building a Psychological World for LLM-Based Agents
- Accelerating A/B-Tests with Counterfactual Estimation: Reducing Variance through Policy Overlap
- Data and trained models for "Empirical Evidence of Large Language Model's Influence on Human Spoken Communication"
- Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
- HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
- Discrete Action Space as a Prerequisite for GRPO Convergence in Small-Model Continuous Control
- Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
- Align AI to Dynamic Human-AI Workflows
- Reliability-Aware LLM Alignment from Inconsistent Human Feedback
- Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models
- One Student, Many Teachers: Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context
- On the Limits of Support-Preserving Alignment and Bounded Filtering
- CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents
- TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
- Normalized Rewards for Preference Optimization
- Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
- From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
- Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework
- Large Language Models Hack Rewards, and Society
- Geometry-Guided Constraint Learning for LLM Safety Classification
- Pretraining Language Models on Historical Text
- Robust Critics: Defending LLMs Against Multi-Turn Attacks
- FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation
- Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
- Preference Tuning as Spectral Update Reorganization
- Hallucinations Undermine Trust; Metacognition is a Way Forward
- S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
- Response drift across frontier large language models
- RE-AD: Real-Time Requirement Adherence for Data Labeling
- The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs
- Machine understanding
- Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs
- Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor
- Deterministic Replay for AI Agent Systems
- PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
- Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
- Brainrot: Deskilling and Addiction are Overlooked AI Risks
- "AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs
- Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- Some Large Language Models Exhibit Consistent Risk Attitudes
- Removing Sandbagging in LLMs by Training with Weak Supervision
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- AI Can Learn Scientific Taste
- Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation
- Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
- Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
- Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
- Discovering Differences in Strategic Behavior Between Humans and LLMs
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- The Illusion of Insight in Reasoning Models
- Self-Distillation Enables Continual Learning
- How Human is AI? Examining the Impact of Emotional Prompts on Artificial and Human and Responsiveness
- Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models
- Legal Alignment for Safe and Ethical AI
- Extracting books from production language models
- No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
- Distributional AGI Safety
- Reasoning Models Ace the CFA Exams
- Epistemological Fault Lines Between Human and Artificial Intelligence
- Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025
- Echoing: Identity Failures when LLM Agents Talk to Each Other
- Are Large Language Models Sensitive to the Motives Behind Communication?
- SAM 3D: 3Dfy Anything in Images
- Kimi Linear: An Expressive, Efficient Attention Architecture
- A Pragmatic View of AI Personhood
- Everyone prefers human writers, including AI
- LLaDA-MoE: A Sparse MoE Diffusion Language Model
- Video models are zero-shot learners and reasoners
- Layer-0 Suppressors Ground Hallucination Inevitability: A Mechanistic Account of How Transformers Trade Factuality for Hedging
- Understanding Reinforcement Learning for Model Training, and future directions with GRAPE
- An Economy of AI Agents
- K2-Think: A Parameter-Efficient Reasoning System
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
- On the Theoretical Limitations of Embedding-Based Retrieval
- Measuring Scalar Constructs in Social Science with LLMs
- Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
- Large Language Models Do Not Simulate Human Psychology
- Training language models to be warm and empathetic makes them less reliable and more sycophantic
- Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
- Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
- Relative Entropy Pathwise Policy Optimization
- Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
- LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
- Cognitive models can reveal interpretable value trade-offs in language models
- Who Does What in Deep Learning? Multidimensional Game-Theoretic Attribution of Function of Neural Units
- LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
- Mercury: Ultra-Fast Language Models Based on Diffusion
- InfoFlood: Jailbreaking Large Language Models with Information Overload
- Self-Adapting Language Models
- Unsupervised Elicitation of Language Models
- Reinforcement Pre-Training
- Securing AI Agents with Information-Flow Control
- A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
- Breaking the Performance Ceiling in Reinforcement Learning requires Inference Strategies
- Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods
- RLSR: Reinforcement Learning from Self Reward
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- Absolute Zero: Reinforced Self-play Reasoning with Zero Data
- El Agente: An autonomous agent for quantum chemistry
- Do Language Models Know Who Did What to Whom?
- Coral Protocol: Open Infrastructure Connecting The Internet of Agents
- Base Models Beat Aligned Models at Randomness and Creativity
- Kongzi: A Historical Large Language Model with Fact Enhancement
- AssistanceZero: Scalably Solving Assistance Games
- Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
- VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
- Not All Data Are Unlearned Equally
- Inference-Time Scaling for Generalist Reward Modeling
- A matter of principle? AI alignment as the fair treatment of claims
- Training large language models on narrow tasks can lead to broad misalignment
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
- Large Language Diffusion Models
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
- Rejected Dialects: Biases Against African American Language in Reward Models
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Re-evaluating Theory of Mind evaluation in large language models
- Why human-AI relationships need socioaffective alignment
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Rethinking Early Stopping: Refine, Then Calibrate
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- A Toolbox for Improving Evolutionary Prompt Search
- The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
- From Accuracy to Impact: The Impact-Driven AI Framework (IDAIF) for Aligning Engineering Architecture with Theory of Change
- The Future of NLP may not be at NLP Conferences: Scholarly Migration Patterns in Natural Language Processing
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories
- Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
- Alpha-R1: Alpha Screening with LLM Reasoning via Reinforcement Learning
- Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
- Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
- C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs
- Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
- MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
- CORE: A Unified Cascaded Ordinal Relevance Estimation Framework for E-commerce Search
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
- EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion
- REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
- The Reward Model Selection Crisis in Personalized Alignment
- Enhancing knowledge graph interactions: A comprehensive Text-to-Cypher pipeline with large language models
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
- APO: Alpha-Divergence Preference Optimization
- Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
- Solipsistic Superintelligence is Unlikely to be Cooperative
- LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
- Eliciting Behaviors in Multi-Turn Conversations
- Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
- ReDiF: Reinforced Distillation for Few Step Diffusion
- Harnessing Large Language Models for Biomedical Named Entity Recognition
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- DICE: Discrete Interpretable Comparative Evaluation with Probabilistic Scoring for Retrieval-Augmented Generation
- M2G-Eval: Enhancing and Evaluating Multi-granularity Multilingual Code Generation
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- Learning When Not to Attend Globally
- Role-Based Fault Tolerance System for LLM RL Post-Training
- DarkPatterns-LLM: A Multi-Layer Benchmark for Detecting Manipulative and Harmful AI Behavior
- Pointer-Augmented Autoregressive Generation of Patent Claims with Joint Topology and Content Decoding
- HalluMat: Detecting Hallucinations in LLM-Generated Materials Science Content Through Multi-Stage Verification
- The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
- Learning from Negative Examples: Why Warning-Framed Training Data Teaches What It Warns Against
- Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
- Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
- Unbiased Visual Reasoning with Controlled Visual Inputs
- SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
- Expert-Grounded Automatic Prompt Engineering for Extracting Lattice Constants of High-Entropy Alloys from Scientific Publications using Large Language Models
- ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning
- Understanding Tone-Dependent Inference Cost in Large Language Models
- Intelligence Without Integrity: Why Capable LLMs May Undermine Reliability
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning
- LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction
- From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
- Scale Weight Decay and Train Better
- Resource-sensitive but language-blind: Community size and not grammatical complexity better predicts the accuracy of Large Language Models in a novel Wug Test
- CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models
- What do Reward Models Memorize?
- Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes
- Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning
- Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models
- Online Policy Evaluation for MDPs with Dynamic UBSR Measures
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- Moral Hazard in Multi-Agent Language Models
- Security and Privacy in Agentic AI: Grand Challenges and Future Directions
- Reward Guided Decoding for Generative Recommendation
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
- Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
- Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
- Inverse RL Helps Align AI by Imitating Humans
- Learning from 53.6K Real-World Developer Edits of AI-Generated Code
- Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
- In-Context Learning as Implicit Policy Gradient
- Traceable LLM Reasoning for Fake-Order Fraud Detection
- Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining
- Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias
- AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning
- Visual Token Compression Enhances Robustness of MLLMs
- MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation
- Retrieval-based and Fine-tuned LLM Approaches for Industrial Asset Health Monitoring and Decision Support
- How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
- PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF
- STAIF: A Stage-wise Optimization for Complex Instruction Following
- ARdena: Scenario-driven control of real-time LLM agents
- DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training
- Group Preference Collapse in Personalized Multimodal Large Language Models
- RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection
- Codifying the Judge: Scalable Evaluation via Program Distillation
- Latent Space Probing for Adult Content Detection in Video Generative Models
- Context Sensitivity Improves Human-Machine Visual Alignment
- The Cartesian Cut in Agentic AI
- Can an Actor-Critic Optimization Framework Improve Analog Design?
- LanteRn: Latent Visual Structured Reasoning
- SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
- Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
- Responsible Intelligence in Practice: A Fairness Audit of Open Large Language Models for Library Reference Services
- El Agente Estructural: An Artificially Intelligent Molecular Editor
- Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis
- SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams
- Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- A Comedy of Estimators: On KL Regularization in RL Training of LLMs
- Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures
- GoldenFuzz: Generative Golden Reference Hardware Fuzzing
- RLLaVA: An RL-central Framework for Language and Vision Assistants
- Beyond Context: Large Language Models Failure to Grasp Users Intent
- Semi-Supervised Learning for Large Language Models Safety and Content Moderation
- Artificial or Just Artful? Do LLMs Bend the Rules in Programming?
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
- DIAL: Direct Iterative Adversarial Learning for Realistic Multi-Turn Dialogue Simulation
- Generalization of RLVR Using Causal Reasoning as a Testbed
- FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
- Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
- BRIDGE: Budget-aware Reasoning via Intermediate Distillation with Guided Examples
- Offline Safe Policy Optimization From Heterogeneous Feedback
- Learning to Reason in LLMs by Expectation Maximization
- Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
- Counterfactual LLM-based Framework for Measuring Rhetorical Style
- Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
- Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation
- Learning General Policies with Policy Gradient Methods
- Humanlike AI Design Increases Anthropomorphism but Yields Divergent Outcomes on Engagement and Trust Globally
- QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation
- AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards
- Explicit and Non-asymptotic Query Complexities of Rank-Based Zeroth-order Algorithm on Stochastic Smooth Functions
- Population-Evolve: a Parallel Sampling and Evolutionary Method for LLM Math Reasoning
- Efficient Personalization of Generative Models via Optimal Experimental Design
- Recontextualization Mitigates Specification Gaming without Modifying the Specification
- ORPR: An OR-Guided Pretrain-then-Reinforce Learning Model for Inventory Management
- Online Robust Reinforcement Learning with General Function Approximation
- FASTRIC: Prompt Specification Language for Verifiable LLM Interactions
- Auto-Prompting with Retrieval Guidance for Frame Detection in Logistics
- MEEA: Mere Exposure Effect-Driven Confrontational Optimization for LLM Jailbreaking
- SafeMed-R1: Adversarial Reinforcement Learning for Generalizable and Robust Medical Reasoning in Vision-Language Models
- Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement
- RMLer: Synthesizing Novel Objects across Diverse Categories via Reinforcement Mixing Learning
- HARBOR: Holistic Adaptive Risk assessment model for BehaviORal healthcare
- Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
- LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction
- LLMs on Drugs: Language Models Are Few-Shot Consumers
- A Multi-agent Text2SQL Framework using Small Language Models and Execution Feedback
- Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
- Software Vulnerability Management in the Era of Artificial Intelligence: An Industry Perspective
- CrystalFormer-CSP: Thinking Fast and Slow for Crystal Structure Prediction
- Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation
- Stable and Efficient Single-Rollout RL for Multimodal Reasoning
- Trustworthy and Explainable Deep Reinforcement Learning for Safe and Energy-Efficient Process Control: A Use Case in Industrial Compressed Air Systems
- ReGal: A First Look at PPO-based Legal AI for Judgment Prediction and Summarization in India
- Adversarial Robustness of Vision in Open Foundation Models
- ShareChat: A Dataset of Chatbot Conversations in the Wild
- AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
- Subjective Question Generation and Answer Evaluation using NLP
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
- Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
- Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
- Synthelite: Chemist-aligned and feasibility-aware synthesis planning with LLMs
- Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
- Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
- MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
- PAACE: A Plan-Aware Automated Agent Context Engineering Framework
- Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
- Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues
- Stepwise Think-Critique: A Unified Framework for Robust and Interpretable LLM Reasoning
- Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
- Evaluating LLMs for Zeolite Synthesis Event Extraction (ZSEE): A Systematic Analysis of Prompting Strategies
- MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
- The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
- Spectral Representation-based Reinforcement Learning
- Model Agnostic Preference Optimization for Medical Image Segmentation
- DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
- Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
- Entropy-Reservoir Bregman Projection: An Information-Geometric Unification of Model Collapse
- Learning to Extract Context for Context-Aware LLM Inference
- IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
- Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
- Inflation Attitudes of Large Language Models
- Georeferencing complex relative locality descriptions with large language models
- History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
- Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
- Understanding and Improving Hyperbolic Deep Reinforcement Learning
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- A First-Order Logic-Based Alternative to Reward Models in RLHF
- A Multifaceted Analysis of Social Biases in Large Language Models
- OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
- Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
- Explainable reinforcement learning from human feedback to improve alignment
- CAPE: Capability Achievement via Policy Execution
- Towards Effective Model Editing for LLM Personalization
- Towards Interactive Intelligence for Digital Humans
- A Scientific Reasoning Model for Organic Synthesis Procedure Generation
- State-Dependent Refusal and Learned Incapacity in RLHF-Aligned Language Models
- MedCEG: Reinforcing Verifiable Medical Reasoning with Critical Evidence Graph
- MiniLingua: A Small Open-Source LLM for European Languages
- Differentiable Evolutionary Reinforcement Learning
- AIR: Post-training Data Selection for Reasoning via Attention Head Influence
- Post-Training and Test-Time Scaling of Generative Agent Behavior Models for Interactive Autonomous Driving
- Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
- Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
- SpeakRL: Synergizing Reasoning, Speaking, and Acting in Language Models with Reinforcement Learning
- Socratic Students: Teaching Language Models to Learn by Asking Questions
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
- Revisiting the Reliability of Language Models in Instruction-Following
- What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation
- Anchoring Values in Temporal and Group Dimensions for Flow Matching Model Alignment
- How Prompts Move Language Model Behavior: Frames, Salience, and Construal as Semantic Control
- HyperEdit: Unlocking Instruction-based Text Editing in LLMs via Hypernetworks
- Adaptive Detector-Verifier Framework for Zero-Shot Polyp Detection in Open-World Settings
- Unified Control for Inference-Time Guidance of Denoising Diffusion Models
- The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior
- Taxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Models
- Rethinking Expert Trajectory Utilization in LLM Post-training
- Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization
- Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
- Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
- RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
- ReactorFold: Generative discovery of nuclear reactor cores via emergent physical reasoning
- Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
- Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems
- MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents
- Your plan may succeed, but what about failure? Investigating how people use ChatGPT for long-term life task planning
- TAO-Net: Two-stage Adaptive OOD Classification Network for Fine-grained Encrypted Traffic Classification
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems
- LLM-Auction: Generative Auction towards LLM-Native Advertising
- Multi-dimensional Preference Alignment by Conditioning Reward Itself
- Exploring LLMs for Scientific Information Extraction Using The SciEx Framework
- A Unified Generative-Predictive Framework for Deterministic Inverse Design
- KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question Answering
- MedBioRAG: Semantic Search and Retrieval-Augmented Generation with Large Language Models for Medical and Biological QA
- Rethinking Chain-of-Thought Reasoning for Videos
- System Report for CCL25-Eval Task 10: Prompt-Driven Large Language Model Merge for Fine-Grained Chinese Hate Speech Detection
- Building Reasonable Inference for Vision-Language Models in Blind Image Quality Assessment
- Chasing Shadows: Pitfalls in LLM Security Research
- RouteRAG: Efficient Retrieval-Augmented Generation from Text and Graph via Reinforcement Learning
- Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment
- Encoder-Free Knowledge-Graph Reasoning with LLMs via Hyperdimensional Path Retrieval
- A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
- Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
- Gradient-Informed Monte Carlo Fine-Tuning of Diffusion Models for Low-Thrust Trajectory Design
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- Uncertainty-Aware Data-Efficient AI: An Information-Theoretic Perspective
- rSIM: Incentivizing Reasoning Capabilities of LLMs via Reinforced Strategy Injection
- A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
- Universal Adversarial Suffixes for Language Models Using Reinforcement Learning with Calibrated Reward
- Large Language Models for Education and Research: An Empirical and User Survey-based Analysis
- Provable Long-Range Benefits of Next-Token Prediction
- Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
- SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
- Depth-Wise Activation Steering for Honest Language Models
- MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
- ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
- ContextualSHAP : Enhancing SHAP Explanations Through Contextual Language Generation
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
- Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
- Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels
- VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
- The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
- Parent-Guided Semantic Reward Model (PGSRM): Embedding-Based Reward Functions for Reinforcement Learning of Transformer Language Models
- Rhea: Role-aware Heuristic Episodic Attention for Conversational LLMs
- Agency at the Interface: Distinguishing Teleological from Structural Self-Organization via Internal Coarse-Graining and Downward Causation
- Cognitive Control Architecture (CCA): A Lifecycle Supervision Framework for Robustly Aligned AI Agents
- Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
- Personalized Image Descriptions from Attention Sequences
- PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
- RLAX: Large-Scale, Distributed Reinforcement Learning for Large Language Models on TPUs
- When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
- Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models
- Auto-exploration for online reinforcement learning
- Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge
- Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning
- ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language Models
- A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
- ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
- ARCANE: A Multi-Agent Framework for Interpretable and Configurable Alignment
- PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
- LLM Harms: A Taxonomy and Discussion
- Intrinsically Interpretable Attention via Sparse Post-Training
- The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
- ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
- Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
- The Road of Adaptive AI for Precision in Cybersecurity
- SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
- Dynamic Alignment for Collective Agency: Toward a Scalable Self-Improving Framework for Open-Ended LLM Alignment
- Value Gradient Guidance for Flow Matching Alignment
- LMSpell: Neural Spell Checking for Low-Resource Languages
- AI & Human Co-Improvement for Safer Co-Superintelligence
- STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
- STELLA: Guiding Large Language Models for Time Series Forecasting with Semantic Abstractions
- SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge Graphs
- Reflection-Satisfaction Tradeoff: Investigating Impact of Reflection on Student Engagement with AI-Generated Programming Hints
- YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
- RLHFSpec: Breaking the Efficiency Bottleneck in RLHF Training via Adaptive Drafting
- Towards an AI Fluid Scientist: LLM-Powered Scientific Discovery in Experimental Fluid Mechanics
- Efficient Reinforcement Learning with Semantic and Token Entropy for LLM Reasoning
- Cross-Task Benchmarking and Evaluation of General-Purpose and Code-Specific Large Language Models
- Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective
- ADAPT: Learning Task Mixtures for Budget-Constrained Instruction Tuning
- VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
- ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
- Data-regularized Reinforcement Learning for Diffusion Models at Scale
- Towards better dense rewards in Reinforcement Learning Applications
- Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order
- Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
- Enhancing Instruction-Following Capabilities in Seq2Seq Models: DoLA Adaptations for T5
- In-Context Representation Hijacking
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- FFTrainer: Fast Failover in Large-Language Model Training with Almost-Free State Management
- Us-vs-Them bias in Large Language Models
- Tipping the Dominos: Topology-Aware Multi-Hop Attacks on LLM-Based Multi-Agent Systems
- Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value
- From static to adaptive: immune memory-based jailbreak detection for large language models
- Idea-Gated Transformers: Enforcing Semantic Coherence via Differentiable Vocabulary Pruning
- PretrainZero: Reinforcement Active Pretraining
- MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking
- Invasive Context Engineering to Control Large Language Models
- Hypothesis Testing for Generalized Thurstone Models
- Network Self-Configuration based on Fine-Tuned Small Language Models
- SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment
- Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models
- Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics Shifts
- Guided Self-Evolving LLMs with Minimal Human Supervision
- OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
- promptolution: A Unified, Modular Framework for Prompt Optimization
- From monoliths to modules: Decomposing transducers for efficient world modelling
- Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning
- AlignSAE: Concept-Aligned Sparse Autoencoders
- Artemis: Structured Visual Reasoning for Perception Policy Learning
- Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
- GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment
- Agentic Policy Optimization via Instruction-Policy Co-Evolution
- Rectifying LLM Thought from Lens of Optimization
- Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
- Exploring Human Perceptions of AI Responses: Insights from a Mixed-Methods Study on Risk Mitigation in Generative Models
- OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
- CauSight: Learning to Supersense for Visual Causal Discovery
- How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
- Zero-Overhead Introspection for Adaptive Test-Time Compute
- AI-Enabled grading with near-domain data for scaling feedback with human-level accuracy
- PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards
- Financial Instruction Following Evaluation (FIFE)
- S2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
- When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
- On The Finetuning of MLIPs Through the Lens of Iterated Maps With BPTT
- Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
- Towards Active Synthetic Data Generation for Finetuning Language Models
- ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
- Clinical-R1: Empowering Large Language Models for Faithful and Comprehensive Reasoning with Clinical Objective Relative Policy Optimization
- Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
- UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
- SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
- When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
- Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation
- Aligning Probabilistic Beliefs under Informative Missingness: LLM Steerability in Clinical Reasoning
- EduEval: A Hierarchical Cognitive Benchmark for Evaluating Large Language Models in Chinese Education
- Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
- Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
- Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards
- Instruction Tuning of Large Language Models for Tabular Data Generation-in One Day
- Listwise Preference Optimization with Element-wise Confusions for Aspect Sentiment Quad Prediction
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
- OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
- Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
- Co-Evolving Agents: Learning from Failures as Hard Negatives
- Optimizing NetGPT via Routing-Based Synergy and Reinforcement Learning
- Tacit Bidder-Side Collusion: Artificial Intelligence in Dynamic Auctions
- TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
- PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and Fuzz Optimization
- Decomposed Trust: Exploring Privacy, Adversarial Robustness, Fairness, and Ethics of Low-Rank LLMs
- MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
- Video Generation Models Are Good Latent Reward Models
- Optimizing Life Sciences Agents in Real-Time using Reinforcement Learning
- Bootstrapping LLMs via Preference-Based Policy Optimization
- The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training
- GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs
- Can Finetuing LLMs on Small Human Samples Increase Heterogeneity, Alignment, and Belief-Action Coherence?
- Self-Guided Defense: Adaptive Safety Alignment for Reasoning Models via Synthesized Guidelines
- TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models
- OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
- ICPO: Intrinsic Confidence-Driven Group Relative Preference Optimization for Efficient Reinforcement Learning
- Escaping the Verifier: Learning to Reason via Demonstrations
- Towards Audio Token Compression in Large Audio Language Models
- Reinforcement Learning for Latent-Space Thinking in LLMs
- Reinforcing Action Policies by Prophesying
- MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models
- On Evaluating LLM Alignment by Evaluating LLMs as Judges
- DesignPref: Capturing Personal Preferences in Visual Design Generation
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
- Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
- Thinking in 360°: Humanoid Visual Search in the Wild
- AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models
- CREward: A Type-Specific Creativity Reward Model
- HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
- Differential Smoothing Mitigates Sharpening and Improves LLM Reasoning
- Profile-LLM: Dynamic Profile Optimization for Realistic Personality Expression in LLMs
- CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
- Learning to Clean: Reinforcement Learning for Noisy Label Correction
- SOMBRL: Scalable and Optimistic Model-Based RL
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
- Advances and Challenges in Solar Flare Prediction: A Review
- Schema Matching on Graph: Iterative Graph Exploration for Efficient and Explainable Data Integration
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
- EAGER: Edge-Aligned LLM Defense for Robust, Efficient, and Accurate Cybersecurity Question Answering
- Large Language Models as Search Engines: Societal Challenges
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Beyond Reward Margin: Rethinking and Resolving Likelihood Displacement in Diffusion Models via Video Generation
- OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
- Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
- Accelerating Reinforcement Learning via Error-Related Human Brain Signals
- Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
- FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
- Optimizing LLM Code Suggestions: Feedback-Driven Timing with Lightweight State Bounds
- RoguePrompt: Dual-Layer Ciphering for Self-Reconstruction to Circumvent LLM Moderation
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation
- Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation
- Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
- Test-Time Preference Optimization for Image Restoration
- Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
- TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- Exploring Weak-to-Strong Generalization for CLIP-based Classification
- Building Domain-Specific Small Language Models via Guided Data Generation
- Efficient Inference Using Large Language Models with Limited Human Data: Fine-Tuning then Rectification
- Curvature-Aware Safety Restoration In LLMs Fine-Tuning
- SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
- Alignment Faking - the Train -> Deploy Asymmetry: Through a Game-Theoretic Lens with Bayesian-Stackelberg Equilibria
- MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
- Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
- QuickLAP: Quick Language-Action Preference Learning for Semi-Autonomous Agents
- Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently
- MultiGA: Leveraging Multi-Source Seeding in Genetic Algorithms
- Mesh RAG: Retrieval Augmentation for Autoregressive Mesh Generation
- The PLLuM Instruction Corpus
- PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
- Why Do Language Model Agents Whistleblow?
- Supervised Fine Tuning of Large Language Models for Domain Specific Knowledge Graph Construction:A Case Study on Hunan's Historical Celebrities
- FIRM: Federated In-client Regularized Multi-objective Alignment for Large Language Models
- Evaluating Adversarial Vulnerabilities in Modern Large Language Models
- Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
- Cross-cultural value alignment frameworks for responsible AI governance: Evidence from China-West comparative analysis
- MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
- SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation
- Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
- Personalized Reward Modeling for Text-to-Image Generation
- Towards Unified Vision Language Models for Forest Ecological Analysis in Earth Observation
- SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning
- "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
- Pass@k Metric for RLVR: A Diagnostic Tool of Exploration, But Not an Objective
- PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
- Multi-Agent Collaborative Reward Design for Enhancing Reasoning in Reinforcement Learning
- Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models
- A Mathematical Framework for Custom Reward Functions in Job Application Evaluation using Reinforcement Learning
- Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
- Exploring Syntropic Frameworks in AI Alignment: A Philosophical Investigation
- Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Reflexive Evidence-Based Multimodal Learning for Clean Energy Transitions: Causal Insights on Cooking Fuel Access, Urbanization, and Carbon Emissions
- Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
- Efficiency Will Not Lead to Sustainable Reasoning AI
- GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
- EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
- SafeRBench: A Comprehensive Benchmark for Safety Assessment in Large Reasoning Models
- Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
- Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
- GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
- GPS: General Per-Sample Prompter
- Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior
- Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning
- Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distribution
- Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
- Just Asking Questions: Doing Our Own Research on Conspiratorial Ideation by Generative AI Chatbots
- TaoSearchEmb: A Multi-Objective Reinforcement Learning Framework for Dense Retrieval in Taobao Search
- VLMs Guided Interpretable Decision Making for Autonomous Driving
- Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
- Generalist Foundation Models Are Not Clinical Enough for Hospital Operations
- ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
- Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
- Evaluating the Ability of Large Language Models to Identify Adherence to CONSORT Reporting Guidelines in Randomized Controlled Trials: A Methodological Evaluation Study
- BeDiscovER: The Benchmark of Discourse Understanding in the Era of Reasoning Language Models
- STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
- Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting
- The Future of Food: How Artificial Intelligence is Transforming Food Manufacturing
- RLHF May Not Reflect Genuine Preferences
- The Alignment Game: A Theory of Long-Horizon Alignment Through Recursive Curation
- LLM Reinforcement in Context
- Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation
- Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
- RAGSmith: A Framework for Finding the Optimal Composition of Retrieval-Augmented Generation Methods Across Datasets
- Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing
- Group-Aware Reinforcement Learning for Output Diversity in Large Language Models
- Detecting LLM-Assisted Academic Dishonesty using Keystroke Dynamics
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
- Tailored Primitive Initialization is the Secret Key to Reinforcement Learning
- From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
- Maximizing the efficiency of human feedback in AI alignment: a comparative analysis
- Mitigating Length Bias in RLHF through a Causal Lens
- AlignTree: Efficient Defense Against LLM Jailbreak Attacks
- Multi-Value Alignment for LLMs via Value Decorrelation and Extrapolation
- Rethinking Deep Alignment Through The Lens Of Incomplete Learning
- MetaGDPO: Alleviating Catastrophic Forgetting with Metacognitive Knowledge through Group Direct Preference Optimization
- EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
- On the Entropy Calibration of Language Models
- PIRA: Preference-Oriented Instruction-Tuned Reward Models with Dual Aggregation
- A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
- Context-Emotion Aware Therapeutic Dialogue Generation: A Multi-component Reinforcement Learning Approach to Language Models for Mental Health Support
- Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping
- Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
- Bytes of a Feather: Personality and Opinion Alignment Effects in Human-AI Interaction
- EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
- VIDEOP2R: Video Understanding from Perception to Reasoning
- Dynamic Temperature Scheduler for Knowledge Distillation
- From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
- Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
- AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
- Persona-Aware Alignment Framework for Personalized Dialogue Generation
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
- In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback
- Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
- Black-Box On-Policy Distillation of Large Language Models
- Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
- LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
- SlideBot: A Multi-Agent Framework for Generating Informative, Reliable, Multi-Modal Presentations
- NSL-MT: Linguistically Informed Negative Samples for Efficient Machine Translation in Low-Resource Languages
- Where does an LLM begin computing an instruction?
- AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
- C3TG: Conflict-aware, Composite, and Collaborative Controlled Text Generation
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
- RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation
- Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning
- Automatic Minds: Cognitive Parallels Between Hypnotic States and Large Language Model Processing
- Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
- HalluClean: A Unified Framework to Combat Hallucinations in LLMs
- Who Gets the Reward, Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
- The Path Not Taken: RLVR Provably Learns Off the Principals
- Reinforcement Learning Control of Quantum Error Correction
- How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
- Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
- DPRM: A Dual Implicit Process Reward Model in Multi-Hop Question Answering
- AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
- Prompt Tuning for Natural Language to SQL with Embedding Fine-Tuning and RAG
- Alignment-Aware Quantization for LLM Safety
- DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
- SERL: Self-Examining Reinforcement Learning on Open-Domain
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
- Majority Rules: LLM Ensemble is a Winning Approach for Content Categorization
- Distributionally Robust Online Markov Game with Linear Function Approximation
- PC-Diffusion: Aligning Diffusion Models with Human Preferences via Preference Classifier
- Judging by the Rules: Compliance-Aligned Framework for Modern Slavery Statement Monitoring
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
- Cortex AISQL: A Production SQL Engine for Unstructured Data
- A Self-Improving Architecture for Dynamic Safety in Large Language Models
- On the Creativity of AI Agents
- What can LLMs tell us about the mechanisms behind polarity illusions in humans? Experiments across model scales and training steps
- RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services
- Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
- HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection
- EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Textual Self-attention Network: Test-Time Preference Optimization through Textual Gradient-based Attention
- SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
- MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
- RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
- FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language Models
- Convergence of Actor-Critic Learning for Mean Field Games and Mean Field Control in Continuous Spaces
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- EASE: Practical and Efficient Safety Alignment for Small Language Models
- What Makes Reasoning Invalid: Echo Reflection Mitigation for Large Language Models
- Adaptive Regularization for Large-Scale Sparse Feature Embedding Models
- Synthetic Data-Driven Prompt Tuning for Financial QA over Tables and Documents
- OpenVLN: Open-world Aerial Vision-Language Navigation
- KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
- Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data
- Kunlun Anomaly Troubleshooter: Enabling Kernel-Level Anomaly Detection and Causal Reasoning for Large Model Distributed Inference
- L2T-Hyena: Enhancing State-Space Models with an Adaptive Learn-to-Teach Framework
- Lived Experience in Dialogue: Co-designing Personalization in Large Language Models to Support Youth Mental Well-being
- A Representation Sharpening Framework for Zero Shot Dense Retrieval
- RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
- Optimizing Diversity and Quality through Base-Aligned Model Collaboration
- Steering Language Models with Weight Arithmetic
- PreResQ-R1: Towards Fine-Grained Rank-and-Score Reinforcement Learning for Visual Quality Assessment via Preference-Response Disentangled Policy Optimization
- Reasoning on Time-Series for Financial Technical Analysis
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- Building Specialized Software-Assistant ChatBot with Graph-Based Retrieval-Augmented Generation
- Reflective Personalization Optimization: A Post-hoc Rewriting Framework for Black-Box Large Language Models
- Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
- Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning
- You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
- Explore Data Left Behind in Reinforcement Learning for Reasoning Language Models
- CPO: Condition Preference Optimization for Controllable Image Generation
- Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization
- Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
- Black-Box Guardrail Reverse-engineering Attack
- Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs
- Advancing Equitable AI: Evaluating Cultural Expressiveness in LLMs for Latin American Contexts
- MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
- RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
- SynQuE: Estimating Synthetic Dataset Quality Without Annotations
- Test-Time Adaptation for LLM Agents via Environment Interaction
- GRAD: Graph-Retrieved Adaptive Decoding for Hallucination Mitigation
- STARS: Segment-level Token Alignment with Rejection Sampling in Large Language Models
- Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology
- Learning Without Critics? Revisiting GRPO in Classical Reinforcement Learning Environments
- LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
- DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents
- Control Barrier Function for Aligning Large Language Models
- COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
- Silenced Biases: The Dark Side LLMs Learned to Refuse
- A Collaborative Reasoning Framework for Anomaly Diagnostics in Underwater Robotics
- Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMs
- Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge
- Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions
- PublicAgent: Multi-Agent Design Principles From an LLM-Based Open Data Analysis Framework
- AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
- Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning
- Extending RLVR to Open-Ended Tasks via Verifiable Multiple-Choice Reformulation
- Understanding New-Knowledge-Induced Factual Hallucinations in LLMs: Analysis and Interpretation
- The Realignment Problem: When Right becomes Wrong in LLMs
- Directional-Clamp PPO
- AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda
- Can Conversational AI Counsel for Change? A Theory-Driven Approach to Supporting Dietary Intentions in Ambivalent Individuals
- An Automated Framework for Strategy Discovery, Retrieval, and Evolution in LLM Jailbreak Attacks
- SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
- LLMs as Judges: Toward The Automatic Review of GSN-compliant Assurance Cases
- Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
- Inference-Time Personalized Alignment with a Few User Preference Queries
- Automated Reward Design for Gran Turismo
- DL4Proteins Jupyter Notebooks Teach how to use Artificial Intelligence for Biomolecular Structure Prediction and Design
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
- Random Initialization of Gated Sparse Adapters
- RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
- 3EED: Ground Everything Everywhere in 3D
- Efficient Test-Time Retrieval Augmented Generation
- DPO-F+: Aligning Code Repair Feedback with Developers' Preferences
- Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learning
- IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
- Do Math Reasoning LLMs Help Predict the Impact of Public Transit Events?
- A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
- Ariadne: A Controllable Framework for Probing and Extending VLM Reasoning Boundaries
- DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching
- Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- Reversal Invariance in Autoregressive Language Models
- Reimagining Safety Alignment with An Image
- Diverse Human Value Alignment for Large Language Models via Ethical Reasoning
- VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision
- Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents
- Disrupting Networks: Amplifying Social Dissensus via Opinion Perturbation and Large Language Models
- Characterizing Selective Refusal Bias in Large Language Models
- Addressing Longstanding Challenges in Cognitive Science with Language Models
- BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
- Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
- Closing the Expression Gap in LLM Instructions via Socratic Questioning
- MolChord: Structure-Sequence Alignment for Protein-Guided Drug Design
- Reasoning Up the Instruction Ladder for Controllable Language Models
- Kad: A Framework for Proxy-based Test-time Alignment with Knapsack Approximation Deferral
- FlowMesh: A Service Fabric for Composable LLM Workflows
- SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- Value Drifts: Tracing Value Alignment During LLM Post-Training
- LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- Data-Efficient RLVR via Off-Policy Influence Guidance
- OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
- BOTS: A Unified Framework for Bayesian Online Task Selection in LLM Reinforcement Finetuning
- Offline Clustering of Preference Learning with Active-data Augmentation
- Graph-Enhanced Policy Optimization in LLM Agent Training
- Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
- Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
- Similarity-Distance-Magnitude Language Models
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
- One Model to Critique Them All: Rewarding Agentic Tool-Use via Efficient Reasoning
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
- LLMBisect: Breaking Barriers in Bug Bisection with A Comparative Analysis Pipeline
- Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
- Approximating Human Preferences Using a Multi-Judge Learned System
- The Information-Theoretic Imperative: Compression and the Epistemic Foundations of Intelligence
- ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
- EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis
- Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves
- Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs
- Not ready for the bench: LLM legal interpretation is unstable and out of step with human judgments
- Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines
- MemSFT: Mitigating Alignment Tax with an External Parametric Memory
- NormWorlds-CF: Solver-Verified Counterfactual Normative Reasoning with Metamorphic-Relation GRPO
- BioDisclose: An Actionability-Aware Benchmark for Biomedical Safety under Adversarial Elicitation
- RecoReward: Recommender-Guided Multimodal Description Generation for Recommendation
- DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
- Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation
- CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
- Deep neural networks and humans both benefit from compositional language structure
- Fairness and Bias in Algorithmic Hiring: A Multidisciplinary Survey
- LLM-Augmented Computational Phenotyping of Long Covid
- Multi-Decoder OneRec: Controllable Generative Retrieval for Multi-Objective Industrial Recommendation
- HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models
- Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
- Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
- Scientific Knowledge Discovery in the Age of Large Language Models
- Constitutional Midtraining: Content Presence Drives Alignment Gains
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
- Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation
- The Innate Economic Preferences of Language Models
- Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback
- SciMON: Scientific Inspiration Machines Optimized for Novelty
- Simulating Subjects: The Promise and Peril of Artificial Intelligence Stand-Ins for Social Agents and Interactions
- RAPID: An Efficient Reinforcement Learning Algorithm for Small Language Models
- Large Language Models for Software Engineering: A Systematic Literature Review
- Large language models propagate race-based medicine
- Enhancing Hate Speech Detection with Fine-Tuned Large Language Models Requires High-Quality Data
- Beware of botshit: How to manage the epistemic risks of generative chatbots
- Aligning Pedagogy with Generative AI: An Approach to Customizing Educational GPTs
- Developing Students’ Statistical Expertise Through Writing in the Age of AI
- Take Goodhart Seriously: Principled Limit on General-Purpose AI Optimization
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- Reward Models are Metrics in a Trench Coat
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
- Can Large Language Models Transform Computational Social Science?
- OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment
- Latent-IM: Latent Interaction Management for Speech LLMs
- DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models
- Safety from Honesty in a Disinterested AI Predictor
- The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
- Vision-Language Models Suppress Female Representations Under Ambiguous Input
- ChatGPT and me: First-time and experienced users’ perceptions of ChatGPT’s communicative ability as a dialogue partner
- Learning by teaching with <scp>ChatGPT</scp> : The effect of teachable <scp>ChatGPT</scp> agent on programming education
- Large language models for biomedicine: foundations, opportunities, challenges, and best practices
- Process Matters more than Output for Distinguishing Humans from Machines
- Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
- How Can Recommender Systems Benefit from Large Language Models: A Survey
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
- Fine-Tuning GPT-5 for GPU Kernel Generation
- Urban-R1: Reinforced MLLMs Mitigate Geospatial Biases for Urban General Intelligence
- LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
- Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
- RL makes MLLMs see better than SFT
- Large language models encode clinical knowledge
- On the Use of Large Language Models for Qualitative Synthesis
- Agents at Risk: How Users Unwittingly Undermine LLM Safety
- Ministral 3
- Investigating the Validity Evidence of Automated Scoring Methods for Divergent Thinking Assessments
- LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs
- Dense and Diverse Goal Coverage in Multi Goal Reinforcement Learning
- PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMs
- Sharpness-Guided Group Relative Policy Optimization via Probability Shaping
- Ideology-Based LLMs for Content Moderation
- Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
- DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA
- Model-Document Protocol for AI Search
- LISTEN to Your Preferences: An LLM Framework for Multi-Objective Selection
- A Survey on Unlearning in Large Language Models
- DEBATE: A Large-Scale Benchmark for Role-Playing LLM Agents in Multi-Agent, Long-Form Debates
- Learning-Based vs Human-Derived Congestion Control: An In-Depth Experimental Study
- Reasoning-Aware GRPO using Process Mining
- FT-ARM: Fine-Tuned Agentic Reflection Multimodal Language Model for Pressure Ulcer Severity Classification with Reasoning
- Greedy Sampling Is Provably Efficient for RLHF
- SPICE: Self-Play In Corpus Environments Improves Reasoning
- OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
- LoRA-DA: Data-Aware Initialization for Low-Rank Adaptation via Asymptotic Analysis
- Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning
- Towards Transparent Reasoning: What Drives Faithfulness in Large Language Models?
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- Semi-Supervised Preference Optimization with Limited Feedback
- MASPRM: Multi-Agent System Process Reward Model
- The Sign Estimator: LLM Alignment in the Face of Choice Heterogeneity
- World Simulation with Video Foundation Models for Physical AI
- Fortytwo: Swarm Inference with Peer-Ranked Consensus
- Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation
- Towards AI as Colleagues: Multi-Agent System Improves Structured Professional Ideation
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- GIFT: Group-relative Implicit Fine Tuning Integrates GRPO with DPO and UNA
- Assessing the Relational Abilities of Large Language Models and Large Reasoning Models
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- On the Impossibility of Retrain Equivalence in Machine Unlearning
- Does GenAI Rewrite How We Write? An Empirical Study on Two-Million Preprints
- The Burden of Interactive Alignment with Inconsistent Preferences
- Temporal Blindness in Multi-Turn LLM Agents: Misaligned Tool Use vs. Human Time Perception
- Debiasing Reward Models by Representation Learning with Guarantees
- Think Twice: Branch-and-Rethink Reasoning Reward Model
- Lightweight Robust Direct Preference Optimization
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- Rethinking Error: “Hallucinations” and Epistemological Indifference
- Neural language models as content analysis tools in psychology
- Larger and more instructable language models become less reliable
- POPI: Personalizing LLMs via Optimized Natural Language Preference Inference
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
- Large language model-based task planning for service robots: A review
- Adaptive Blockwise Search: Inference-Time Alignment for Large Language Models
- Code Aesthetics with Agentic Reward Feedback
- Smaller Models, Smarter Rewards: A Two-Sided Approach to Process and Outcome Rewards
- Can Language Models Compose Skills In-Context?
- MGFRec: Towards Reinforced Reasoning Recommendation with Multiple Groundings and Feedback
- Offline Preference Optimization via Maximum Marginal Likelihood Estimation
- Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust
- Retracing the Past: LLMs Emit Training Data When They Get Lost
- Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models
- FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
- Sentra-Guard: A Multilingual Human-AI Framework for Real-Time Defense Against Adversarial LLM Jailbreaks
- Aligning Diffusion Language Models via Unpaired Preference Optimization
- Scalable Oversight via Partitioned Human Supervision
- Frustratingly Easy Task-aware Pruning for Large Language Models
- Agent-GSPO: Communication-Efficient Multi-Agent Systems via Group Sequence Policy Optimization
- Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
- Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
- A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
- GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
- PACR: Progressively Ascending Confidence Reward for LLM Reasoning
- You Don't Need Prompt Engineering Anymore: The Prompting Inversion
- DETECT: Determining Ease and Textual Clarity of German Text Simplifications
- OlaMind: Towards Human-Like and Hallucination-Safe Customer Service for Retrieval-Augmented Dialogue
- Controllable Mathematical Reasoning via Self-Optimizing Thought Vectors
- Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
- Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- α-LoRA: Effective Fine-Tuning via Base Model Rescaling
- Weak-to-Strong Generalization under Distribution Shifts
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
- Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
- Social Simulations with Large Language Model Risk Utopian Illusion
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generation
- Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
- Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- Epipolar Geometry Improves Video Generation Models
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 Samples
- Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
- Language Ranker: A Lightweight Ranking framework for LLM Decoding
- Robust Preference Alignment via Directional Neighborhood Consensus
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- BoundRL: Efficient Structured Text Segmentation through Reinforced Boundary Generation
- Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning
- Vox-Evaluator: Enhancing Stability and Fidelity for Zero-shot TTS with A Multi-Level Evaluator
- Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
- No-Regret Thompson Sampling for Finite-Horizon Markov Decision Processes with Gaussian Processes
- An Empirical Study of Sample Selection Strategies for Large Language Model Repair
- KL-Regularized Reinforcement Learning is Designed to Mode Collapse
- Black Box Absorption: LLMs Undermining Innovative Ideas
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows
- Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
- g-DPO: Scalable Preference Optimization for Protein Language Models
- Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
- Enhancing visual-LLM for construction site safety compliance via prompt engineering and Bi-stage retrieval-augmented generation
- Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
- Data-Centric Lessons To Improve Speech-Language Pretraining
- FairGRPO: Fair Reinforcement Learning for Equitable Clinical Reasoning
- Review of Tools for Zero-Code LLM Based Application Development
- SynCast: Synergizing Contradictions in Precipitation Nowcasting via Diffusion Sequential Preference Optimization
- PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models
- HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- No Compute Left Behind: Rethinking Reasoning and Sampling with Masked Diffusion Models
- RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
- The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- QKCV Attention: Enhancing Time Series Forecasting with Static Categorical Embeddings for Both Lightweight and Pre-trained Foundation Models
- Search Self-play: Pushing the Frontier of Agent Capability without Supervision
- KAT-Coder Technical Report
- Verifiable Accuracy and Abstention Rewards in Curriculum RL to Alleviate Lost-in-Conversation
- HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
- Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
- Large language models in medicine
- Extracting alignment data in open models
- Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
- Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
- StarBench: A Turn-Based RPG Benchmark for Agentic Multimodal Decision-Making and Information Seeking
- Chain-of-Conceptual-Thought Elicits Daily Conversation in Large Language Models
- AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library
- ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
- ADPO: Anchored Direct Preference Optimization
- ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control
- IF-VidCap: Can Video Caption Models Follow Instructions?
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
- Heterogeneous Adversarial Play in Interactive Environments
- DP2O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution
- Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning
- Mapping Post-Training Forgetting in Language Models at Scale
- Planned Diffusion
- Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
- Assessing Monotone Dependence: Area Under the Curve Meets Rank Correlation
- Unbiased Gradient Low-Rank Projection
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- A Principle of Targeted Intervention for Multi-Agent Reinforcement Learning
- Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
- A Mimamsa Inspired Framework For Instruction Sequencing In AI Agents
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
- Zero‐ and few‐shot prompting of generative large language models provides weak assessment of risk of bias in clinical trials
- Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning
- Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
- The Atomic Instruction Gap: Instruction-Tuned LLMs Struggle with Simple, Self-Contained Directives
- Strengthening LLMs for Tabular Prediction with Structural Priors
- Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
- Intent-Driven LLM Ensemble Planning for Flexible Multi-Robot Disassembly: Demonstration on EV Batteries
- Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models
- Fine-tuning Flow Matching Generative Models with Intermediate Feedback
- Rewarding the Journey, Not Just the Destination: A Composite Path and Answer Self-Scoring Reward Mechanism for Test-Time Reinforcement Learning
- JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
- Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
- OG-Rank: Learning to Rank Fast and Slow with Uncertainty and Reward-Trend Guided Adaptive Exploration
- Annotation-Efficient Universal Honesty Alignment
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models
- SafeSearch: Do Not Trade Safety for Utility in LLM Search Agents
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
- QuanBench: Benchmarking Quantum Code Generation with Large Language Models
- Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
- Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
- Dual-Weighted Reinforcement Learning for Generative Preference Modeling
- ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
- InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training
- Voting with the Graph: Stable RLAIF via Topological Consistency Maximization
- The Road Less Traveled: Enhancing Exploration in LLMs via Sequential Sampling
- MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
- Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
- Stochastic Optimization with Random Search
- MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
- STABLE: Gated Continual Learning for Large Language Models
- Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning
- Continual Learning via Sparse Memory Finetuning
- DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
- Reasoning with Sampling: Your Base Model is Smarter Than You Think
- Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
- Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
- Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
- Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
- Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
- Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback
- Inference-Time Search using Side Information for Diffusion-based Image Reconstruction
- Your Next Token Prediction: A Multilingual Benchmark for Personalized Response Generation
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
- Stop-RAG: Value-Based Retrieval Control for Iterative RAG
- Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm
- Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space
- Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation
- Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
- Where to Search: Measure the Prior-Structured Search Space of LLM Agents
- Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning
- Budget-aware Test-time Scaling via Discriminative Verification
- Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
- Training LLM Agents to Empower Humans
- Confidence as a Reward: Transforming LLMs into Reward Models
- M2PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation
- Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling
- PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
- From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails
- Beyond Imitation: Recovering Dense Rewards from Demonstrations
- Improved Robustness of Deep Reinforcement Learning for Control of Time-Varying Systems by Bounded Extremum Seeking
- A11YN: aligning LLMs for accessible web UI code generation
- The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
- ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
- SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
- Automated Network Protocol Testing with LLM Agents
- Deliberate Lab: A Platform for Real-Time Human-AI Social Experiments
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
- Repairing Reward Functions with Human Feedback to Mitigate Reward Hacking
- On the Role of Preference Variance in Preference Optimization
- A Survey on Evaluation of Large Language Models
- How Well Can Preference Optimization Generalize Under Noisy Feedback?
- Data-Model Co-Evolution: Growing Test Sets to Refine LLM Behavior
- From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models
- Expert or not? assessing data quality in offline reinforcement learning
- COSTAR-A: A prompting framework for enhancing Large Language Model performance on Point-of-View questions
- Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
- Laminar: A Scalable Asynchronous RL Post-Training Framework
- VISaGE: Understanding Visual Generics and Exceptions
- Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
- Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
- PromptFlow: Training Prompts Like Neural Networks
- Reinforced Preference Optimization for Recommendation
- ResearStudio: A Human-Intervenable Framework for Building Controllable Deep-Research Agents
- Self-Verifying Reflection Helps Transformers with CoT Reasoning
- Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editing
- Locket: Robust Feature-Locking Technique for Language Models
- Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback
- Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
- Hierarchical Alignment: Surgical Fine-Tuning via Functional Layer Specialization in Large Language Models
- CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
- Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
- ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution
- Pruning Cannot Hurt Robustness: Certified Trade-offs in Reinforcement Learning
- EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
- Don't Walk the Line: Boundary Guidance for Filtered Generation
- Analyzing and Internalizing Complex Policy Documents for LLM Agents
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
- Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
- DocReward: A Document Reward Model for Structuring and Stylizing
- Vision-LLMs for Spatiotemporal Traffic Forecasting
- Exploring and Leveraging Class Vectors for Classifier Editing
- AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
- Automating Structural Engineering Workflows with Large Language Model Agents
- APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
- Rediscovering Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
- Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
- BanglaMATH : A Bangla benchmark dataset for testing LLM mathematical reasoning at grades 6, 7, and 8
- Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
- Does LLM Focus on the Right Words? Mitigating Context Bias in LLM-based Recommenders
- AMiD: Knowledge Distillation for LLMs with α-mixture Assistant Distribution
- Direct Multi-Token Decoding
- Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
- Learning Dynamics of VLM Finetuning
- UpSafe^∘C: Upcycling for Controllable Safety in Large Language Models
- Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
- Controllable Generative Trajectory Prediction via Weak Preference Alignment
- Hierarchical Optimization via LLM-Guided Objective Evolution for Mobility-on-Demand Systems
- DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
- Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
- MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
- PrediQL: Automated Testing of GraphQL APIs with LLMs
- VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
- OpusAnimation: Code-Based Dynamic Chart Generation
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
- Reasoning-Enhanced Large Language Models for Molecular Property Prediction
- You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
- PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under Subspace Calibration
- CompassNav: Steering From Path Imitation To Decision Understanding In Navigation
- Breaking the Likelihood Trap: Consistent Generative Recommendation with Graph-structured Model
- Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models
- Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
- Beyond the limitation of a single query: Train your LLM for query expansion with Reinforcement Learning
- RIPRAG: Hack a Black-box Retrieval-Augmented Generation Question-Answering System with Reinforcement Learning
- Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
- A-IPO: Adaptive Intent-driven Preference Optimization
- Artificial intelligence as a surrogate brain: Bridging neural dynamical models and data
- Contemplative Superalignment
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
- Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
- Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces
- Advancing conversational diagnostic AI with multimodal reasoning
- On the Role of Domain Experts in Creating Effective Tutoring Systems
- Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
- IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
- Signals: Trajectory Sampling and Triage for Agentic Interactions
- Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
- An Alternative Trajectory for Generative AI
- Can RL Improve Generalization of LLM Agents? An Empirical Study
- Don't Throw Away Your Pretrained Model
- Token Is All You Price
- SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models
- ConDABench: Interactive Evaluation of Language Models for Data Analysis
- Agentic Systems in Radiology: Design, Applications, Evaluation, and Challenges
- Safety Game: Balancing Safe and Informative Conversations with Blackbox Agentic AI using LP Solvers
- DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
- GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis
- DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
- AdaPM: a Partial Momentum Algorithm for LLM Training
- Student Development Agent: Risk-free Simulation for Evaluating AIED Innovations
- Users as Annotators: LLM Preference Learning from Comparison Mode
- Leading the Follower: Learning Persuasive Agents in Social Deduction Games
- Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
- SHERLOCK: Towards Dynamic Knowledge Adaptation in LLM-enhanced E-commerce Risk Management
- SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
- Score-Based Density Estimation from Pairwise Comparisons
- Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Markov Likelihood
- Active Model Selection for Large Language Models
- KORMo: Korean Open Reasoning Model for Everyone
- Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models
- VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search
- Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
- From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP
- Opponent Shaping in LLM Agents
- LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens
- Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
- Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
- AILoRA: Function-Aware Asymmetric Initialization for Low-Rank Adaptation of Large Language Models
- LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?
- Contrastive Weak-to-strong Generalization
- Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
- MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
- Energy-Driven Steering: Reducing False Refusals in Large Language Models
- Stop DDoS Attacking the Research Community with AI-Generated Survey Papers
- Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models
- GCPO: When Contrast Fails, Go Gold
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
- Next-Generation LLM for UAV: From Natural Language to Autonomous Flight
- Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization
- Position: Privacy Is Not Just Memorization!
- AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
- VisualDAN: Exposing Vulnerabilities in VLMs with Visual-Driven DAN Commands
- On the optimization dynamics of RLVR: Gradient gap and step size thresholds
- Drift No More? Context Equilibria in Multi-Turn LLM Interactions
- MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces
- Self-Improving LLM Agents at Test-Time
- Reinforcing Diffusion Models by Direct Group Preference Optimization
- LiveThinking: Enabling Real-Time Efficient Reasoning for AI-Powered Livestreaming via Reinforcement Learning
- Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy Optimization
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- An Adaptive Multi Agent Bitcoin Trading System
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
- TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
- Prepared mind, fast response: A temporal decoupling framework for adaptive knowledge orchestration in open-domain dialogue
- Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
- Post-Norm can Resharpen Attention
- From Data to Rewards: a Bilevel Optimization Perspective on Maximum Likelihood Estimation
- Phase Diagram of Dropout for Two-Layer Neural Networks in the Mean-Field Regime
- MAPRO: Recasting Multi-Agent Prompt Optimization as Maximum a Posteriori Inference
- On the Convergence of Moral Self-Correction in Large Language Models
- Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- Reasoning for Hierarchical Text Classification: The Case of Patents
- AI for Abolition? A Participatory Design Approach
- LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish
- Prompt Optimization Across Multiple Agents for Representing Diverse Human Populations
- Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
- LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling
- Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-Oz
- AMAS: Adaptively Determining Communication Topology for LLM-based Multi-Agent System
- Experiential Reinforcement Learning
- Authenticated Workflows: A Systems Approach to Protecting Agentic AI
- CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
- Generating Meaning: Active Inference and the Scope and Limits of Passive AI
- Predictive Preference Learning from Human Interventions
- DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning
- Do LLMs Know They Are Being Tested? Evaluation Awareness and Incentive-Sensitive Failures in GPT-OSS-20B
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
- Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-based Customer Support
- ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
- Aligning Large Language Models via Fully Self-Synthetic Data
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
- Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation
- Optimal Stopping vs Best-of-N for Inference Time Optimization
- Online Rubrics Elicitation from Pairwise Comparisons
- Reasoning by Exploration: A Unified Approach to Retrieval and Generation over Graphs
- POME: Post Optimization Model Edit via Muon-style Projection
- Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
- EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
- LLM Bias Detection and Mitigation through the Lens of Desired Distributions
- Taxonomy of User Needs and Actions
- The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
- Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
- EARL: Efficient Agentic Reinforcement Learning Systems for Large Language Models
- Prompt reinforcing for long-term planning of large language models
- Optimizing for Persuasion Improves LLM Generalization: Evidence from Quality-Diversity Evolution of Debate Strategies
- EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
- Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
- Primal-Dual Direct Preference Optimization for Constrained LLM Alignment
- On the Role of Difficult Prompts in Self-Play Preference Optimization
- Vul-R2: A Reasoning LLM for Automated Vulnerability Repair
- Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
- When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
- Classical AI vs. LLMs for Decision-Maker Alignment in Health Insurance Choices
- MADIAVE: Multi-Agent Debate for Implicit Attribute Value Extraction
- Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique
- Prototype-Based Dynamic Steering for Large Language Models
- Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
- The Answer Lies Within: Self-Derived Rewards Enable Explainable Relation Extraction
- Bloom: Designing for LLM-Augmented Behavior Change Interactions
- AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering
- InvThink: Premortem Reasoning for Safer Language Models
- Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment
- Chrysalis: A Unified System for Comparing Active Teaching and Passive Learning with AI Agents in Education
- Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
- RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
- From Noisy Traces to Stable Gradients: Bias-Variance Optimized Preference Optimization for Aligning Large Reasoning Models
- TeachLM: Post-Training LLMs for Education Using Authentic Learning Data
- Alignment Tipping Process: How Self-Evolution Pushes LLM Agents Off the Rails
- Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning
- Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
- FT-MDT: Extracting Decision Trees from Medical Texts via a Novel Low-rank Adaptation Method
- EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents
- FedSRD: Sparsify-Reconstruct-Decompose for Communication-Efficient Federated Large Language Models Fine-Tuning
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
- Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
- VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
- DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
- Good Intentions Beyond ACL: Who Does NLP for Social Good, and Where?
- Proactive defense against LLM Jailbreak
- COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
- Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
- GRACE: Generative Representation Learning via Contrastive Policy Optimization
- VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
- LLM Based Bayesian Optimization for Prompt Search
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- Self-supervised diffusion model fine-tuning for costate initialization using Markov chain Monte Carlo
- Spatiotemporal Forecasting as Planning: A Model-Based Reinforcement Learning Approach with Generative World Models
- Reflection Before Action: Designing a Framework for Quantifying Thought Patterns for Increased Self-awareness in Personal Decision Making
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment frm Heterogeneous Rewards
- Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
- Learning from All: Concept Alignment for Autonomous Distillation from Multiple Drifting MLLMs
- Fine-Grained GRPO for Precise Preference Alignment in Flow Models
- Large Language Models Hallucination: A Comprehensive Survey
- Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
- Increasing LLM response trustworthiness using voting ensembles
- Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
- Activation Steering with a Feedback Controller
- JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
- Equipping Retrieval-Augmented Large Language Models with Document Structure Awareness
- AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
- MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
- RESTRAIN: From Spurious Votes to Signals -- Self-Driven RL with Self-Penalization
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- How Catastrophic is Your LLM? Certifying Risk in Conversation
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
- TROLL: Trust Regions improve Reinforcement Learning for Large Language Models
- OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
- TreePrompt: Leveraging Hierarchical Few-Shot Example Selection for Improved English-Persian and English-German Translation
- Group Policy Gradient
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
- Decoupling Task-Solving and Output Formatting in LLM Generation
- Generalized Fitted Q-Iteration with Clustered Data
- Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
- Triplet-Structured Knowledge Integration for Multi-Turn Medical Reasoning
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
- A Granular Study of Safety Pretraining under Model Abliteration
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
- Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Self-Improvement in Multimodal Large Language Models: A Survey
- Smart-GRPO: Smartly Sampling Noise for Efficient RL of Flow-Matching Models
- Fine-Tuning on Noisy Instructions: Effects on Generalization and Performance
- Fine-Tuning Jailbreaks under Highly Constrained Black-Box Settings: A Three-Pronged Approach
- Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
- GRAD: Generative Retrieval-Aligned Demonstration Sampler for Efficient Few-Shot Reasoning
- Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?
- mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
- Strategic Fusion of Vision Language Models: Shapley-Credited Context-Aware Dawid-Skene for Multi-Label Tasks in Autonomous Driving
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Multi-Actor Multi-Critic Deep Deterministic Reinforcement Learning with a Novel Q-Ensemble Method
- GEM: A Gym for Agentic LLMs
- Uncovering the Computational Ingredients of Human-Like Representations in LLMs
- It Takes Two: Your GRPO Is Secretly DPO
- From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
- Inclusive Easy-to-Read Generation for Individuals with Cognitive Impairments
- ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
- On Predictability of Reinforcement Learning Dynamics for Large Language Models
- Rethinking Reward Models for Multi-Domain Test-Time Scaling
- Structuring Reasoning for Complex Rules Beyond Flat Representations
- A Call to Action for a Secure-by-Design Generative AI Paradigm
- Train on Validation (ToV): Fast data selection with applications to fine-tuning
- Retrieval-Augmented Framework for LLM-Based Clinical Decision Support
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- Making, not Taking, the Best of N
- BroRL: Scaling Reinforcement Learning via Broadened Exploration
- Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning across the Model Capability Continuum
- Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
- Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
- Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
- Prompt Curriculum Learning for Efficient LLM Post-Training
- Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- Judging with Confidence: Calibrating Autoraters to Preference Distributions
- Debunk the Myth of SFT Generalization
- GRPO-λ: Credit Assignment improves LLM Reasoning
- Thinkquel: A Model Dedicated to Text-to-dbt Using Synthetic Data and a Span-Aware Objective
- PrimeX: A Dataset of Worldview, Opinion, and Explanation
- Which Rewards Matter? Reward Selection for Reinforcement Learning under Limited Feedback
- Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
- Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning
- SafeBehavior: Simulating Human-Like Multistage Reasoning to Mitigate Jailbreak Attacks in Large Language Models
- One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient
- Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
- Alignment-Aware Decoding
- DyFlow: Dynamic Workflow Framework for Agentic Reasoning
- RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation
- RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- Efficient On-Policy Reinforcement Learning via Exploration of Sparse Parameter Space
- Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Supporting Creative Ownership through Deep Learning-Based Music Variation
- Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
- Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking
- Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
- OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- MuPlon: Multi-Path Causal Optimization for Claim Verification through Controlling Confounding
- The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- PCPO: Proportionate Credit Policy Optimization for Aligning Image Generation Models
- Learning to Reason as Action Abstractions with Scalable Mid-Training RL
- Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
- RFG: Test-Time Scaling for Diffusion Large Language Model Reasoning with Reward-Free Guidance
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- IRIS: Intrinsic Reward Image Synthesis
- Aligning Multilingual Reasoning with Verifiable Semantics from a High-Resource Expert Model
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
- Toxicity in Online Platforms and AI Systems: A Survey of Needs, Challenges, Mitigations, and Future Directions
- A Method for Quantifying Human Risk and a Blueprint for LLM Integration
- Fingerprinting LLMs via Prompt Injection
- Polychromic Objectives for Reinforcement Learning
- Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
- UniAPL: A Unified Adversarial Preference Learning Framework for Instruct-Following
- The Era of Real-World Human Interaction: RL from User Conversations
- Rethinking Entropy Regularization in Large Reasoning Models
- CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- SecInfer: Preventing Prompt Injection via Inference-time Scaling
- Learning Distinguishable Representations in Deep Q-Networks for Linear Transfer
- Retro*: Optimizing LLMs for Reasoning-Intensive Document Retrieval
- T-POP: Test-Time Personalization with Online Preference Feedback
- Reference-Free Rating of LLM Responses via Latent Information
- CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task
- Multilingual Text-to-SQL: Benchmarking the Limits of Language Models with Collaborative Language Agents
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
- PEARL: Performance-Enhanced Aggregated Representation Learning
- DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
- Graph Optimization Foundation Model: Tokenizing Graph via A Language-Model Paradigm
- Prompt and Parameter Co-Optimization for Large Language Models
- Humanline: Online Alignment as Perceptual Loss
- Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
- RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
- World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
- The problem of alignment
- Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
- Interactive Groupwise Comparison for Reinforcement Learning from Human Feedback
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Agentar-Scale-SQL: Advancing Text-to-SQL through Orchestrated Test-Time Scaling
- Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
- Reinforcement Mid-Training
- Preference-Based Dynamic Ranking Structure Recognition
- Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE
- EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
- ADAPT: Lightweight, Long-Range Machine Learning Force Fields Without Graphs
- Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training
- MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- Reinforcement Learning with Inverse Rewards for World Model Post-training
- Winning the Pruning Gamble: A Unified Approach to Joint Sample and Token Pruning for Efficient Supervised Fine-Tuning
- Rethinking Reward Miscalibration of GRPO in Agentic RL
- Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
- Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
- Anchored Supervised Fine-Tuning
- Why Alignment Must Precede Distillation: A Minimal Working Explanation
- Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs
- Fast Thinking for Large Language Models
- Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
- On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
- Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
- Advancing Multi-agent Traffic Simulation via R1-Style Reinforcement Fine-Tuning
- EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance
- Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models
- Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
- Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
- MedCritical: Enhancing Medical Reasoning in Small Language Models via Self-Collaborative Correction
- Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
- Decoupling Reasoning and Perception: An LLM-LMM Framework for Faithful Visual Reasoning
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Knowledge distillation through geometry-aware representational alignment
- Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts
- Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs
- General Exploratory Bonus for Optimistic Exploration in RLHF
- Multiplayer Nash Preference Optimization
- DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning
- Risk Profiling and Modulation for LLMs
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
- WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning
- Adaptive Margin RLHF via Preference over Preferences
- Causally-Enhanced Reinforcement Policy Optimization
- Reinforcement Learning-Based Prompt Template Stealing for Text-to-Image Models
- MTRec: Learning to Align with User Preferences via Mental Reward Models
- <scp>ChatGPT</scp> for complex text evaluation tasks
- Voice user interfaces for effortless navigation in medical virtual reality environments
- A shared model-based linguistic space for transmitting our thoughts from brain to brain in natural conversations
- Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
- Optimizing Long-Form Clinical Text Generation with Claim-Based Rewards
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
- Boosting Pointer Analysis With LLM-Enhanced Allocation Function Detection
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- RAPID3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
- Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models
- When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
- Context Parametrization with Compositional Adapters
- Multilingual Vision-Language Models, A Survey
- S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
- The Rogue Scalpel: Activation Steering Compromises LLM Safety
- Goal-Guided Efficient Exploration via Large Language Model in Reinforcement Learning
- Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
- Exposing Hallucinations To Suppress Them: VLMs Representation Editing With Generative Anchors
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
- Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
- Synthetic Dialogue Generation for Interactive Conversational Elicitation & Recommendation (ICER)
- SoK: Potentials and Challenges of Large Language Models for Reverse Engineering
- Can LLMs Solve and Generate Linguistic Olympiad Puzzles?
- FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
- ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
- Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
- AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
- Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
- In-Context Learning can Perform Continual Learning Like Humans
- Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization
- We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
- Compute-Optimal Quantization-Aware Training
- RLP: Reinforcement as a Pretraining Objective
- Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models
- MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning
- DriftLite: Lightweight Drift Control for Inference-Time Scaling of Diffusion Models
- Multi-Objective Reinforcement Learning for Large Language Model Optimization: Visionary Perspective
- Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
- Correct Reasoning Paths Visit Shared Decision Pivots
- Plan2Evolve: LLM Self-Evolution for Improved Planning Capability via Automated Domain Generation
- Preemptive Detection and Steering of LLM Misalignment via Latent Reachability
- Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
- Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
- The role of synthetic data in Multilingual, Multi-cultural AI systems: Lessons from Indic Languages
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
- LLMTrace: A Corpus for Classification and Fine-Grained Localization of AI-Written Text
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
- LogReasoner: Empowering LLMs with Expert-like Coarse-to-Fine Reasoning for Automated Log Analysis
- Physics of Learning: A Lagrangian perspective to different learning paradigms
- CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
- It's Not You, It's Clipping: A Soft Trust-Region via Probability Smoothing for LLM RL
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning
- Difference-Guided Reasoning: A Temporal-Spatial Framework for Large Language Models
- Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
- PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
- Actor-Critic without Actor
- Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
- d2: Improving Reasoning in Diffusion Language Models via Trajectory Likelihood Estimation
- Enhancing Python Programming Education with an AI-Powered Code Helper: Design, Implementation, and Impact
- Complexity-Regularized Proximal Policy Optimization
- Instruction Boundary: Quantifying Biases in LLM Reasoning under Various Coverage
- ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
- Failure Modes of Maximum Entropy RLHF
- OR-Toolformer: Modeling and Solving Operations Research Problems with Tool Augmented Large Language Models
- Federation of Agents: A Semantics-Aware Communication Fabric for Large-Scale Agentic AI
- V-GameGym: Visual Game Generation for Code Large Language Models
- LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generation
- Embodied AI: From LLMs to World Models
- The Knowledge-Behaviour Disconnect in LLM-based Chatbots
- MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
- Future Policy Aware Preference Learning for Mathematical Reasoning
- PolicyPad: Collaborative Prototyping of LLM Policies
- Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation
- Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
- UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
- bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
- ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
- From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
- Training Skills Like Parameters via Self-Supervised Semantic Diffusion
- What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
- RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
- DeepResearch Agent System
- Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
- ACPO: Asymmetric Credit Policy Optimization via Mode-Local Entropy Surrogate
- MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
- FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
- LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware
- Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
- IFHierBench: Hierarchical Instruction Following for Large Language Models
- Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
- Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models
- SDO: Structure-Aware Data Organization for Efficient LLM Post-Training
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
- RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
- MedLLM: An Open Medical Language Model at the Sub-Billion Scale
- ToolRec: Calibrated Preference Alignment for Query Recommendation in On-Device Assistants
- Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
- Divergence Decoding: Training-Free Capability Fusion
- Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
- Machine learning in computational literary studies
- A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing
- DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes
- Implicit Safety Alignment from Crowd Preferences
- Persuading large language models to comply with objectionable requests
- Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
- S-GRPO: Unified Post-Training for Large Vision-Language Models
- Paper Espresso: From Paper Overload to Research Insight
- RELISH: LLM REgression with a Latent Iterative State Head
- GenAI and the Mirage of Personalised Learning for All
- From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
- Game-theory behaviour of large language models: The case of Keynesian beauty contests
- A Theory of Appropriateness That Accounts for Norms of Rationality
- A large language model-based agent for wayfinding: simulation of spatial perception and memory
- Pressure Reveals Character: Behavioural Alignment Evaluation at Depth
- References Improve LLM Alignment in Non-Verifiable Domains
- SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
- Assessing and alleviating state anxiety in large language models
- Security and Privacy Challenges of Large Language Models: A Survey
- Auditing large language models: a three-layered approach
- Assessing political bias in AI systems: a framework for disentangling viewpoint preferences from epistemic integrity
- Bypassing Guardrails: Lessons Learned from Red Teaming ChatGPT
- Toward Human-Centered Explainability: Natural Language Explanations for Anomaly Detection
- GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
- Summary of ChatGPT-Related research and perspective towards the future of large language models
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
- Reinforcement Learning on Pre-Training Data
- Agentic Reinforcement Learning with Implicit Step Rewards
- Speculative Safety-Aware Decoding
- SMITE: Enhancing Fairness in LLMs through Optimal In-Context Example Selection via Dynamic Validation
- Central Limit Theorems for Asynchronous Averaged Q-Learning
- Direct Preference Optimization for Speech Autoregressive Diffusion Models
- Diversity Boosts AI-Generated Text Detection
- MAPO: Mixed Advantage Policy Optimization
- SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer
- Enhancing LLM-Based Social Bot via an Adversarial Learning Framework
- PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generation
- LLMZ+: Contextual Prompt Whitelist Principles for Agentic LLMs
- APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- NGRPO: Negative-enhanced Group Relative Policy Optimization
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- Advances in Large Language Models for Medicine
- GRPO++: Enhancing Dermatological Reasoning under Low Resource Settings
- Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation
- The Narcissus Hypothesis: Descending to the Rung of Illusion
- ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- Automated Knowledge Graph Construction using Large Language Models and Sentence Complexity Modelling
- Exploiting Tree Structure for Credit Assignment in RL Training of LLMs
- Weights-Rotated Preference Optimization for Large Language Models
- LAD-VF: LLM-Automatic Differentiation Enables Fine-Tuning-Free Robot Planning from Formal Methods Feedback
- An Artificial Intelligence Value at Risk Approach: Metrics and Models
- ATLAS: Benchmarking and Adapting LLMs for Global Trade via Harmonized Tariff Code Classification
- A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialogue
- PG-CE: A Progressive Generation Dataset with Constraint Enhancement for Controllable Text Generation
- Understanding Post-Training Structural Changes in Large Language Models
- DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
- nDNA -- the Semantic Helix of Artificial Cognition
- IDfRA: Self-Verification for Iterative Design in Robotic Assembly
- Advancing Speech Understanding in Speech-Aware Language Models with GRPO
- Preference Distillation via Value based Reinforcement Learning
- SemanticGarment: Semantic-Controlled Generation and Editing of 3D Gaussian Garments
- Can GRPO Boost Complex Multimodal Table Understanding?
- LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
- SFT-TA: Supervised Fine-Tuned Agents in Multi-Agent LLMs for Automated Inductive Thematic Analysis
- USB-Rec: An Effective Framework for Improving Conversational Recommendation Capability of Large Language Model
- Improving User Interface Generation Models from Designer Feedback
- Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories
- Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models
- From Uniform to Heterogeneous: Tailoring Policy Optimization to Every Token's Nature
- RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
- Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
- Assessing Classical Machine Learning and Transformer-based Approaches for Detecting AI-Generated Research Text
- The Programmer’s Assistant: Conversational Interaction with a Large Language Model for Software Development
- RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
- Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
- SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning
- The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
- Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation
- Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
- The Alignment Bottleneck
- Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- REFER: Mitigating Bias in Opinion Summarisation via Frequency Framed Prompting
- Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
- The governance & behavioral challenges of generative artificial intelligence’s hypercustomization capabilities
- LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
- Reward Hacking Mitigation using Verifiable Composite Rewards
- How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
- Jamendo-QA: A Large-Scale Music Question Answering Dataset
- Dynamic Classifier-Free Diffusion Guidance via Online Feedback
- BaseReward: A Strong Baseline for Multimodal Reward Model
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
- Real, Fake, or Manipulated? Detecting Machine-Influenced Text
- Self-Improving Embodied Foundation Models
- AutoEdit: Automatic Hyperparameter Tuning for Image Editing
- CARGO: A Framework for Confidence-Aware Routing of Large Language Models
- Rationality Check! Benchmarking the Rationality of Large Language Models
- LLM Jailbreak Detection for (Almost) Free!
- Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
- TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
- Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
- Assessing Historical Structural Oppression Worldwide via Rule-Guided Prompting of Large Language Models
- CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
- Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
- Synthetic bootstrapped pretraining
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- Dual-Actor Fine-Tuning of VLA Models: A Talk-and-Tweak Human-in-the-Loop Approach
- SAIL-VL2 Technical Report
- Learning the natural history of human disease with generative transformers
- An LLM-based multi-agent framework for agile effort estimation
- DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models
- Improving Context Fidelity via Native Retrieval-Augmented Reasoning
- DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
- Programmable Cognitive Bias in Social Agents
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
- SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
- A deep reinforcement learning platform for antibiotic discovery
- Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews
- RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features
- Participatory AI: A Scandinavian Approach to Human-Centered AI
- Don't Change My View: Ideological Bias Auditing in Large Language Models
- Yet Another Watermark for Large Language Models
- Gender-Neutral Rewriting in Italian: Models, Approaches, and Trade-offs
- Shaping Explanations: Semantic Reward Modeling with Encoder-Only Transformers for GRPO
- Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
- Multi-Metric Preference Alignment for Generative Speech Restoration
- SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation
- Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
- Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
- Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
- Active Domain Knowledge Acquisition with 100-Dollar Budget: Enhancing LLMs via Cost-Efficient, Expert-Involved Interaction in Sensitive Domains
- Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
- Scaling Agents via Continual Pre-training
- Human resource management in the age of generative artificial intelligence: Perspectives and research directions on ChatGPT
- Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
- Prompt Commons: Collective Prompting as Governance for Urban AI
- Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
- GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Models
- Exploring Conversational Design Choices in LLMs for Pedagogical Purposes: Socratic and Narrative Approaches for Improving Instructor's Teaching Practice
- MusicSwarm: Biologically Inspired Intelligence for Music Composition
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
- Pluralistic Off-policy Evaluation and Alignment
- Audited Reasoning Refinement: Fine-Tuning Language Models via LLM-Guided Step-Wise Evaluation and Correction
- When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- What Matters in Data for DPO?
- ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
- Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
- Pathological Truth Bias in Vision-Language Models
- Gradient Free Deep Reinforcement Learning With TabPFN
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- Beyond IVR Touch-Tones: Customer Intent Routing using LLMs
- Continually Adding New Languages to Multilingual Language Models
- Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
- A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment
- Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning
- AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
- When Are Two RLHF Objectives the Same?
- Limitations of refinement methods for weak to strong generalization
- The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language Models
- RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems
- SearchInstruct: Enhancing Domain Adaptation via Retrieval-Based Instruction Dataset Creation
- CrunchLLM: Multitask LLMs for Structured Business Reasoning and Outcome Prediction
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Pluralistic Alignment for Healthcare: A Role-Driven Framework
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
- DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillation
- WebSight: A Vision-First Architecture for Robust Web Agents
- VARCO-VISION-2.0 Technical Report
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Multi-Intent Recognition in Dialogue Understanding: A Comparison Between Smaller Open-Source LLMs
- InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- Towards Understanding Visual Grounding in Visual Language Models
- Inpainting-Guided Policy Optimization for Diffusion Large Language Models
- Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration
- Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks
- Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining
- Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
- Latency and Token-Aware Test-Time Compute
- How well can LLMs provide planning feedback in grounded environments?
- Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling
- RewardDance: Reward Scaling in Visual Generation
- X-Teaming Evolutionary M2S: Automated Discovery of Multi-turn to Single-turn Jailbreak Templates
- Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing
- Generative Data Refinement: Just Ask for Better Data
- CM-Align: Consistency-based Multilingual Alignment for Large Language Models
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- Ubiquitous Intelligence Via Wireless Network-Driven LLMs Evolution
- Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- Language Self-Play For Data-Free Training
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward
- Uncovering Scaling Laws for Large Language Models via Inverse Problems
- EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention
- Fuzz4All: Universal Fuzzing with Large Language Models
- Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
- SoK: Security and Privacy of AI Agents for Blockchain
- Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
- UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
- MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
- Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
- Reinforcement Learning Foundations for Deep Research Systems: A Survey
- Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks
- Outcome-based Exploration for LLM Reasoning
- Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation
- A Fragile Number Sense: Probing the Elemental Limits of Numerical Reasoning in LLMs
- The Thinking Therapist: Training Large Language Models to Deliver Acceptance and Commitment Therapy using Supervised Fine-Tuning and Odds Ratio Policy Optimization
- RL Fine-Tuning Heals OOD Forgetting in SFT
- Sovereign AI for 6G: Towards the Future of AI-Native Networks
- Uncovering the Vulnerability of Large Language Models in the Financial Domain via Risk Concealment
- EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
- Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
- IntrEx: A Dataset for Modeling Engagement in Educational Conversations
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
- Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
- BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
- Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL
- Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
- Benchmarking Gender and Political Bias in Large Language Models
- Reverse-Engineered Reasoning for Open-Ended Generation
- Finetuning LLMs for Human Behavior Prediction in Social Science Experiments
- Understanding the Influence of Synthetic Data for Text Embedders
- Coefficients-Preserving Sampling for Reinforcement Learning with Flow Matching
- RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
- Chatbot To Help Patients Understand Their Health
- CC-GSEO-Bench: A Content-Centric Benchmark for Measuring Source Influence in Generative Search Engines
- Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation
- ZhiFangDanTai: Fine-tuning Graph-based Retrieval-Augmented Generation Model for Traditional Chinese Medicine Formula
- Self-Aligned Reward: Towards Effective and Efficient Reasoners
- CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
- LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
- Cloning a Conversational Voice AI Agent from Call Recording Datasets for Telesales
- What-If Analysis of Large Language Models: Explore the Game World Using Proactive Thinking
- Post-training Large Language Models for Diverse High-Quality Responses
- A Lightweight Framework for Trigger-Guided LoRA-Based Self-Adaptation in LLMs
- Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
- Symbolic Graphics Programming with Large Language Models
- Murphys Laws of AI Alignment: Why the Gap Always Wins
- Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment
- Towards a Unified View of Large Language Model Post-Training
- Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
- Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology
- Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- On Aligning Prediction Models with Clinical Experiential Learning: A Prostate Cancer Case Study
- HAMSA: Hijacking Aligned Compact Models via Stealthy Automation
- On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
- NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
- MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
- SPFT-SQL: Enhancing Large Language Model for Text-to-SQL Parsing by Self-Play Fine-Tuning
- SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Measuring How (Not Just Whether) VLMs Build Common Ground
- Beyond Interpretability: Exploring the Comprehensibility of Adaptive Video Streaming through Large Language Models
- AI-in-the-Loop: Privacy Preserving Real-Time Scam Detection and Conversational Scambaiting by Leveraging LLMs and Federated Learning
- SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
- PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
- Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
- Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
- ChatGPT-generated texts show authorship traits that identify them as non-human
- TraceLLM: Security Diagnosis Through Traces and Smart Contracts in Ethereum
- Advancing SLM Tool-Use Capability using Reinforcement Learning
- Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
- Towards Reasoning for PDE Foundation Models: A Reward-Model-Driven Inference-Time-Scaling Algorithm
- Scaling behavior of large language models in emotional safety classification across sizes and tasks
- Omnidirectional Spatial Modeling from Correlated Panoramas
- Generative KI für TA
- The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
- Re-evaluating LLM-based Heuristic Search: A Case Study on the 3D Packing Problem
- Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving
- DCPO: Dynamic Clipping Policy Optimization
- GRAM-R2: Self-Training Generative Foundation Reward Models for Reward Reasoning
- ChatOps for microservice systems: A low-code approach using service composition and large language models
- Abex-rat: Synergizing Abstractive Augmentation and Adversarial Training for Classification of Occupational Accident Reports
- Relative Trajectory Balance is equivalent to Trust-PCL
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency
- Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
- Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
- Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation
- Generative Goal Modeling
- On the Alignment of Large Language Models with Global Human Opinion
- Reinforcement Learning for Machine Learning Engineering Agents
- CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
- The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
- Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMs
- MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
- Confident, Calibrated, or Complicit: Probing the Trade-offs between Safety Alignment and Ideological Bias in Language Models in Detecting Hate Speech
- Political Ideology Shifts in Large Language Models
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
- Neural Models and Language Model Prompting for the Multidimensional Evaluation of Open-Ended Conversations
- RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
- LLM-Assisted Iterative Evolution with Swarm Intelligence Toward SuperBrain
- Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
- Open Data Synthesis For Deep Research
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- SABR: A Stable Adaptive Bitrate Framework Using Behavior Cloning Pretraining and Reinforcement Learning Fine-Tuning
- GIER: Gap-Driven Self-Refinement for Large Language Models
- VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
- Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
- Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
- Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight
- Introduction to the Analysis of Probabilistic Decision-Making Algorithms
- Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
- Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards
- SoK: Exposing the Generation and Detection Gaps in LLM-Generated Phishing Through Examination of Generation Methods, Content Characteristics, and Countermeasures
- Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
- MultiFluxAI Enhancing Platform Engineering with Advanced Agent-Orchestrated Retrieval Systems
- Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- Personality Matters: User Traits Predict LLM Preferences in Multi-Turn Collaborative Tasks
- BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
- InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity
- AI Reasoning Models for Problem Solving in Physics
- Language-Enhanced Mobile Manipulation for Efficient Object Search in Indoor Environments
- Bridging Minds and Machines: Toward an Integration of AI and Cognitive Science
- NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
- Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
- TCIA: A Task-Centric Instruction Augmentation Method for Instruction Finetuning
- Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
- Learning to Generate Unit Test via Adversarial Reinforcement Learning
- Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
- Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant
- Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
- JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
- On the possibility of deep alignment
- SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models
- 11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
- Model Science: getting serious about verification, explanation and control of AI systems
- IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement
- Evaluating Language Model Reasoning about Confidential Information
- CapTune: Adapting Non-Speech Captions With Anchored Generative Models
- HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
- PSO-Merging: Merging Models Based on Particle Swarm Optimization
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?
- Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models
- Safety Alignment Should Be Made More Than Just A Few Attention Heads
- Position: The Pitfalls of Over-Alignment: Overly Caution Health-Related Responses From LLMs are Unethical and Dangerous
- Skill-based Explanations for Serendipitous Course Recommendation
- Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities
- Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
- MotionFlux: Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment
- Learning Game-Playing Agents with Generative Code Optimization
- Towards 6G Intelligence: The Role of Generative AI in Future Wireless Networks
- Do MLLMs Really Understand the Charts?
- Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
- Ensemble Debates with Local Large Language Models for AI Alignment
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- Active Query Selection for Crowd-Based Reinforcement Learning
- From Bits to Boardrooms: A Cutting-Edge Multi-Agent LLM Framework for Business Excellence
- HAEPO: History-Aggregated Exploratory Policy Optimization
- Recycling History: Efficient Recommendations from Contextual Dueling Bandits
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement
- Beyond Quality: Unlocking Diversity in Ad Headline Generation with Large Language Models
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing
- LaQual: A Novel Framework for Automated Evaluation of LLM App Quality
- TrackRec: Iterative Alternating Feedback with Chain-of-Thought via Preference Alignment for Recommendation
- Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience
- Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
- SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis
- SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
- SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
- Transduction is All You Need for Structured Data Workflows
- Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- Open-Universe Assistance Games
- Mapping the Course for Prompt-based Structured Prediction
- Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
- Reinforcement learning entangling operations on spin qubits
- Linear Preference Optimization: Decoupled Gradient Control via Absolute Regularization
- Distribution-Guided Auto-Encoder for User Multimodal Interest Cross Fusion
- In2x at WMT25 Translation Task
- DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization
- Automated Optimization Modeling through Expert-Guided Large Language Model Reasoning
- DEPTH: Hallucination-Free Relation Extraction via Dependency-Aware Sentence Simplification and Two-tiered Hierarchical Refinement
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- BioLORD-2023: semantic textual representations fusing large language models and clinical knowledge graph insights
- Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
- Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper
- Embedding Democratic Values into Social Media AIs via Societal Objective Functions
- Incident Analysis for AI Agents
- Learning from Preferences and Mixed Demonstrations in General Settings
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- AI Testing Should Account for Sophisticated Strategic Behaviour
- LLMind 2.0: Distributed IoT Automation with Natural Language M2M Communication and Lightweight LLM Agents
- CIA+TA Risk Assessment for AI Reasoning Vulnerabilities
- Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
- LM Agents May Fail to Act on Their Own Risk Knowledge
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- MAVIS: Multi-Objective Alignment via Value-Guided Inference-Time Search
- ALIGN: Word Association Learning for Cultural Alignment in Large Language Models
- LENS: Learning to Segment Anything with Unified Reinforced Reasoning
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning
- DPad: Efficient Diffusion Language Models with Suffix Dropout
- Graph Concept Bottleneck Models
- FLAIR: Feedback Learning for Adaptive Information Retrieval
- Stands to Reason: Investigating the Effect of Reasoning on Idiomaticity Detection
- AI Agents for Photonic Integrated Circuit Design Automation
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Hallucinations in medical devices
- Involuntary Jailbreak: On Self-Prompting Attacks
- Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning
- Consiglieres in the Shadow: Understanding the Use of Uncensored Large Language Models in Cybercrimes
- RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
- Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
- MDPO: Overcoming the Training-Inference Divide of Masked Diffusion Language Models
- Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback
- Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position
- The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
- Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
- Mitigating Jailbreaks with Intent-Aware LLMs
- SafeCtrl: Region-Based Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress
- QuarkMed Medical Foundation Model Technical Report
- In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
- Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings
- Learning Wisdom from Errors: Promoting LLM's Continual Relation Learning through Exploiting Error Cases
- Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation
- Controlling Multimodal LLMs via Reward-guided Decoding
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- Preference Models assume Proportional Hazards of Utilities
- Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions
- From Clicks to Preference: A Multi-stage Alignment Framework for Generative Query Suggestion in Conversational System
- Feedback Indicators: The Alignment between Llama and a Teacher in Language Learning
- Fusing Rewards and Preferences in Reinforcement Learning
- HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
- Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
- Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
- FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
- Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
- Hard Examples Are All You Need: Maximizing GRPO Post-Training Under Annotation Budgets
- SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
- Tailoring Diagnostic Modeling to Individual Learners: Personalized Distractor Generation via MCTS-Guided Reasoning Reconstruction
- Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
- Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
- Reinforced Language Models for Sequential Decision Making
- Agentic Design Review System
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
- ReviewRL: Towards Automated Scientific Review with RL
- Artificial Emotion: A Survey of Theories and Debates on Realising Emotion in Artificial Intelligence
- LingVarBench: Benchmarking LLM for Automated Named Entity Recognition in Structured Synthetic Spoken Transcriptions
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development
- Perturbed Public Voices (P2V): A Dataset for Robust Audio Deepfake Detection
- BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
- What are the limits to biomedical research acceleration through general-purpose AI?
- MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
- Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillation
- On Negative-aware Preference Optimization for Recommendation
- The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
- User-centric Subjective Leaderboard by Customizable Reward Modeling
- COMPEER: Controllable Empathetic Reinforcement Reasoning for Emotional Support Conversation
- ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
- Affordances of Sketched Notations for Multimodal UI Design and Development Tools
- Reducing Cognitive Overhead in Tool Use via Multi-Small-Agent Reinforcement Learning
- BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
- Interpretable Reward Model via Sparse Autoencoder
- STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
- A Survey on Training-free Alignment of Large Language Models
- Special-Character Adversarial Attacks on Open-Source Language Model
- Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
- DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
- Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
- Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
- From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
- PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
- ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation
- DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
- Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
- Generating Query-Relevant Document Summaries via Reinforcement Learning
- SAEMark: Multi-bit LLM Watermarking with Inference-Time Scaling
- Reinforcement Learning for Large Model: A Survey
- WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
- \(X\)-evolve: Solution space evolution powered by large language models
- Learning to Align, Aligning to Learn: A Unified Approach for Self-Optimized Alignment
- ThinkTuning: Instilling Cognitive Reflections without Distillation
- From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR
- Careful Queries, Credible Results: Teaching RAG Models Advanced Web Search Tools with Reinforcement Learning
- Large Language Models for Subjective Language Understanding: A Survey
- Grid2Guide: A* Enabled Small Language Model for Indoor Navigation
- Vision-Based Localization and LLM-based Navigation for Indoor Environments
- Data Selection for LLM Alignment Using Fine-Grained Preferences
- Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints
- NeuroDx-LM: A Clinical Large-Scale Model for EEG-based Neurological Disorder Detection
- Pareto Multi-Objective Alignment for Language Models
- Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks for Enhanced Action Understanding
- Think Before You Talk: Enhancing Meaningful Dialogue Generation in Full-Duplex Speech Language Models with Planning-Inspired Text Guidance
- Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback
- A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection
- A Principled Loss Function for Direct Language Model Alignment
- Pref-GUIDE: Continual Policy Learning from Real-Time Human Feedback via Preference-Based Learning
- AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
- Many-Turn Jailbreaking
- Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models
- Bridging Cultural Nuances in Dialogue Agents through Cultural Value Surveys
- HapticLLaMA: A Multimodal Sensory Language Model for Haptic Captioning
- The Fair Game: Auditing & Debiasing AI Algorithms Over Time
- Sample-efficient LLM Optimization with Reset Replay
- LLM Unlearning Without an Expert Curated Dataset
- Towards Integrated Alignment
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
- Position: Intelligent Coding Systems Should Write Programs with Justifications
- A Framework for Inherently Safer AGI through Language-Mediated Active Inference
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
- Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models
- Iterative Learning of Computable Phenotypes for Treatment Resistant Hypertension using Large Language Models
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
- Mixed-Initiative Dialog for Human-Robot Collaborative Manipulation
- The World According to LLMs: How Geographic Origin Influences LLMs' Entity Deduction Capabilities
- Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
- Can Language Models Critique Themselves? Investigating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025
- mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering
- Decision-Making with Deliberation: Meta-reviewing as a Document-grounded Dialogue
- ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
- QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
- AI-assisted JSON Schema Creation and Mapping
- Posterior-GRPO: Rewarding Reasoning Processes in Code Generation
- Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
- SPaRFT: Self-Paced Reinforcement Fine-Tuning for Large Language Models
- Can Large Language Models Integrate Spatial Data? Empirical Insights into Reasoning Strengths and Computational Weaknesses
- Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning
- Echo: Decoupling Inference and Training for Large-Scale RL Alignment on Heterogeneous Swarms
- Root Cause Analysis Training for Healthcare Professionals With AI-Powered Virtual Simulation: A Proof-of-Concept
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- FaST: Feature-aware Sampling and Tuning for Personalized Preference Alignment with Limited Data
- Multi-module GRPO: Composing Policy Gradients and Prompt Optimization for Language Model Programs
- IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
- P-Aligner: Enabling Pre-Alignment of Language Models via Principled Instruction Synthesis
- TRAIL: Joint Inference and Refinement of Knowledge Graphs with Large Language Models
- SimInstruct: A Responsible Tool for Collecting Scaffolding Dialogues Between Experts and LLM-Simulated Novices
- Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
- Method-Based Reasoning for Large Language Models: Extraction, Reuse, and Continuous Improvement
- Large Language Model's Multi-Capability Alignment in Biomedical Domain
- KG-Augmented Executable CoT for Mathematical Coding
- AttriLens-Mol: Attribute Guided Reinforcement Learning for Molecular Property Prediction with Large Language Models
- Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
- GeoSR: Cognitive-Agentic Framework for Probing Geospatial Knowledge Boundaries via Iterative Self-Refinement
- Generative Bid Shading in Real-Time Bidding Advertising
- Sotopia-RL: Reward Design for Social Intelligence
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- DiWA: Diffusion Policy Adaptation with World Models
- LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations
- LaTCoder: Converting Webpage Design to Code with Layout-as-Thought
- MultiRAG: A Knowledge-guided Framework for Mitigating Hallucination in Multi-source Retrieval Augmented Generation
- Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis
- EvaDrive: Evolutionary Adversarial Policy Optimization for End-to-End Autonomous Driving
- Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling
- V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion Models
- GTPO: Stabilizing Group Relative Policy Optimization via Gradient and Entropy Control
- Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
- Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
- Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
- ORVIT: Near-Optimal Online Distributionally Robust Reinforcement Learning
- Survey of Large Language Models in Extended Reality: Technical Paradigms and Application Frontiers
- On the Evaluation of Large Language Models in Multilingual Vulnerability Repair
- PLoRA: Efficient LoRA Hyperparameter Tuning for Large Models
- CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
- An Efficient and Adaptive Next Edit Suggestion Framework with Zero Human Instructions in IDEs
- CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
- Traffic-R1: Reinforced LLMs Bring Human-Like Reasoning to Traffic Signal Control Systems
- CAAD: Context-Aware Adaptive Decoding for Truthful Text Generation
- A Survey on AgentOps: Categorization, Challenges, and Future Directions
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
- MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
- Uni-Layout: Integrating Human Feedback in Unified Layout Generation and Evaluation
- SAMPO-Path: Segmentation Intent-Aligned Preference Optimization for Pathology Foundation Model Segmentation
- TIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language Models
- Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
- A Survey on Data Security in Large Language Models
- MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
- Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
- Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
- Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention
- Censored Sampling for Topology Design: Guiding Diffusion with Human Preferences
- A Theory of Adaptive Scaffolding for LLM-Based Pedagogical Agents
- TeSent: A Benchmark Dataset for Fairness-aware Explainable Sentiment Classification in Telugu
- From Query to Logic: Ontology-Driven Multi-Hop Reasoning in LLMs
- MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
- Adaptive Content Restriction for Large Language Models via Suffix Optimization
- Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
- RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
- Bias Association Discovery Framework for Open-Ended LLM Generations
- Provably Secure Retrieval-Augmented Generation
- A Note on Code Quality Score: LLMs for Maintainable Large Codebases
- Better Call Claude: Can LLMs Detect Changes of Writing Style?
- Activation-Guided Local Editing for Jailbreaking Attacks
- Foundations of Interpretable Models
- Pro2Guard: Proactive Runtime Enforcement of LLM Agent Safety via Probabilistic Model Checking
- Thinking Machines: Mathematical Reasoning in the Age of LLMs
- Dual Collaborative LLMs via Continual Fine-Tuning for Serendipitous Recommendation
- AutoDebias: Automated Framework for Debiasing Text-to-Image Models
- Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
- Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
- MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning
- MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generation
- Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
- Lucy: edgerunning agentic web search on mobile with machine generated task vectors
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
- Calibrated Language Models and How to Find Them with Label Smoothing
- ITDR: An Instruction Tuning Dataset for Enhancing Large Language Models in Recommendations
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
- GraphRAG-R1: Graph Retrieval-Augmented Generation with Process-Constrained Reinforcement Learning
- Causal2Vec: Improving Decoder-only LLMs as Versatile Embedding Models
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
- AutoBridge: Automating Smart Device Integration with Centralized Platform
- BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
- How Far Are AI Scientists from Changing the World?
- T-Detect: Tail-Aware Statistical Normalization for Robust Detection of Adversarial Machine-Generated Text
- MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
- DynaSwarm: Dynamically Graph Structure Selection for LLM-based Multi-agent System
- Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
- Rethinking Evidence Hierarchies in Medical Language Benchmarks: A Critical Evaluation of HealthBench
- Open-Source Agentic Hybrid RAG Framework for Scientific Literature Review
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- Opportunities and Challenges of LLMs in Education: An NLP Perspective
- OFCnetLLM: Large Language Model for Network Monitoring and Alertness
- ShortFT: Diffusion Model Alignment via Shortcut-based Fine-Tuning
- Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
- Heartificial Intelligence: Exploring Empathy in Language Models
- AI alignment [wikipedia]
- AI safety [wikipedia]
- ChatGPT [wikipedia]
- Post-training of large language models [wikipedia]
- Products and applications of OpenAI [wikipedia]
- Reasoning model [wikipedia]
- Reinforcement learning [wikipedia]
- Reinforcement learning from human feedback [wikipedia]
- Paul Christiano [wikipedia]
Discussions
- Gemini lies to user about health info, says it wanted to make him feel better— Though commonly reported, Google doesn't consider it a security problem when models make things up [lemmy, 104 points, 21 comments]
- "Our models are neither fully aligned nor fully safe; they still generate toxic or biased outputs, make up facts, and generate sexual and violent content without explicit prompting." arxiv.org/abs/22 [bsky, 12 points, 1 comments]
- Training language models to follow instructions with human feedback [pdf] [hn, 3 points, 0 comments]
- Training language models to follow instructions with human feedback [hn, 2 points, 0 comments]
- arxiv.org/abs/2203.02155 Instruct GPT [bsky, 0 points, 0 comments]
- Possibly it does, but only because ChatGPT has in its pedigree InstructGPT, which was specifically trained w/ RLHF on prompts that contained instructions to do things? arxiv.org/abs/2203.02155 [bsky, 0 points, 0 comments]
- 15. InstructGPT arxiv.org/abs/2203.02155 Reading papers is one thing. Understanding why each changed the field is what makes you a stronger AI engineer. [bsky, 0 points, 0 comments]
- ... zu InstructGPT (arxiv.org/abs/2203.02155) mit der LLM nach dem ersten Pre-Training das Frage-/Antwort-Format gelernt wird. Fun Fakt: Einige Crowd Sourced Trainingsdaten versuchten dem erst pre-tra [bsky, 0 points, 1 comments]
- 人間からの入力をフィードバックにする手法はここらへん arxiv.org/abs/2203.02155 > Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of th [bsky, 0 points, 0 comments]
Related