Training language models to follow instructions with human feedback
2022/03/04 by Long Ouyang, Ouyang, Long, Jeff Wu +38 · 9 voices · 4,379 citations
Computer Science · #Artificial intelligence #Computer science #Explainable Artificial Intelligence (XAI) #Human–computer interaction #Language model #Machine learning #Natural Language Processing Techniques #Natural language processing #Programming language #Range (aeronautics) #Reinforcement learning #Set (abstract data type) #Simple (philosophy) #Topic Modeling #Training set #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2203.02155
published in arXiv (Cornell University) (Cornell University)
arxiv created 2022/03/04 · openalex publication_date 2022/03/04 · arxiv updated 2022/03/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
Making language models bigger does not inherently make them better at following a user's intent. For example, large language models can generate outputs that are untruthful, toxic, or simply not helpful to the user. In other words, these models are not aligned with their users. In this paper, we show an avenue for aligning language models with user intent on a wide range of tasks by fine-tuning with human feedback. Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback. We call the resulting models InstructGPT. In human evaluations on our prompt distribution, outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3, despite having 100x fewer parameters. Moreover, InstructGPT models show improvements in truthfulness and reductions in toxic output generation while having minimal performance regressions on public NLP datasets. Even though InstructGPT still makes simple mistakes, our results show that fine-tuning with human feedback is a promising direction for aligning language models with human intent.
Cited by
- ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents
- Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
- When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
- Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning
- Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMs
- Unified Static-Dynamic Pruning for Efficient LLM Inference
- Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
- Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Models
- Three-Body Alignment: Aligning Chess Agent with Human Reasoning through Reranked Rationale
- Hint-Guided Diversified Policy Optimization for LLM Reasoning
- ASEval: Automated Trajectory-Level Security Testing for Autonomous Agents
- Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
- Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives
- Improving Large Vision-Language Models' Understanding for Flow Field Data
- From Mind to Machine: The Rise of Manus AI as a Fully Autonomous Digital Agent
- Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
- Learning to Reason for Factuality
- Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability
- Be Consistent! Enhancing Robust Visual Reasoning in LVLMs with Consistency Constraints
- A Unified Moral-Value Dataset for Instruction Tuning
- PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
- Branching Policy Optimization: Sandbox-Native Language Agent Reinforcement Learning
- Safety boundary maintenance in consumer AI systems responding to pediatric health queries: a cross-platform benchmark evaluation under naturalistic and adversarially pressured conditions
- Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner
- PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses
- SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning
- RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
- AlphaRoute: Large Language Models as Semantic Optimizers for Multi-Objective Routing
- PercepCap: Video Captioner with Structured Spatio-Temporal Perception
- Towards Automated Formal Verification of zkEVMs Using LLM-Guided Constraint Synthesis
- Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
- Antigen-specific Antibody Multi-modal Foundation Model for Functional Antibody Design
- Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents
- Sound Probabilistic Safety Bounds for Large Language Models
- SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement
- Verifiable Self-Evolution for Open-Ended Dialogue Skills via Future-Feedback Prediction
- Adopting Reinforcement Learning with Verifiable Rewards for Molecular Generation
- Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
- Post-Training in Time Series Foundation Models: A Unifying Framework
- Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment
- SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking
- REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning
- Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks
- Cognitive Dual-Process Planning for Autonomous Driving with Structured Scene Knowledge and Verifiable Reasoning-Action Consistency
- How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF
- XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs
- Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models
- Task Competence Is Not Instruction Following: Evaluating Instruction-Conflicting Behavior in Small Language Models
- When Reasoning Narrows the Move: Diversity Collapse in LLM Game Play
- Semantic-Aware Data-Aided Channel Estimation with Large Language Models for MIMO Systems
- Rushes: A Human Preference Dataset for Pluralistic Alignment
- Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
- Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels
- Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models
- Rater State Bias in RLHF Preference Data: An Audit Framework
- Dr. Zero: Self-Evolving Search Agents without Training Data
- RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
- LP-SFT: Local-Preserving Supervised Fine-Tuning via Multimodal Entropy Structure
- Falsifiable Release Gates for Self-Improving Systems: Standing Invariants at Scale
- In-Context Learning for Wound Classification with Small Multimodal Language Models
- PlotTwist: A Creative Plot Generation Framework with Small Language Models
- Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs
- DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization
- Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction
- Discovery by Dreaming: Cross-Domain Recombination in Artificial Memory
- Meta-Learning Preferences for Multilingual LLM Alignment
- Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
- LLM-Driven Cross-Paradigm Design for Quantum Optimal Control
- ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples
- Matching Ranks Over Probability Yields Truly Deep Safety Alignment
- After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions
- BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
- MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
- A Geometric Perspective on Stabilizing Value Conflict Resolution
- Signed Rectified Flow: Negativity-Controlled Generation
- LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications
- OR Else: A Differentiable Trust Region for Policy Optimization
- Assisting or resisting patriarchy? a critical discourse analysis of chatgpt’s responses on feminism
- ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
- RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts
- Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
- DeLIVeR: Decomposed Learning for Information-grounded Veracity Recognition via Reinforced Knowledge Graph Exploration
- AI Value Alignment for Evolving Social Norms
- Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
- Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins
- Rationalizing Boltzmann Rationality: An Axiomatic Characterization of Entropy-Regularized Policies
- Distilled Reinforcement Learning for LLM Post-training
- Trace-Based On-Policy Distillation for Masked Diffusion Language Models
- The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
- The Behavioral Credibility Trilemma: When Calibrated Autonomy Becomes Impossible
- Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL
- Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis
- LogicIF: Towards Complex Logic Instruction Following
- Model-Driven Discipline for Multi-Agent LLMs: Requirement-to-Verification Generation of Traceable System Models
- RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
- Diversity-Oriented Fine-Tuning for Uncertainty-Based Hallucination Detection
- Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
- Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets
- Scaling Point-in-Time Language Models
- Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
- TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue
- RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation
- Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
- Decoupled Alignment for Robust Plug-and-Play Adaptation
- Nonuniformity Principle in Human-AI Coworking
- One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models
- Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function
- Auditing Inference-Time Defense Evaluation for Multimodal Large Language Models
- When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
- Relational Preference Encoding in Looped Transformer Internal States
- Understanding Agent-Reactive Bugs at the Model-Harness Boundary: An Empirical Study of LLM Agent Issue Reports
- The CRAFT principles for the responsible use of large language models in policymaking
- EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
- Mask-Aware Policy Gradients for Diffusion Language Models
- Step-Level Preference Learning for Generative Agents in Social Simulations
- DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
- Multi-Turn On-Policy Distillation with Prefix Replay
- Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
- CoTu at EXACT 2026: Neuro-Symbolic Reasoning for Transparent Educational QA
- Scaling Evaluation-time Compute with Reasoning Models as Evaluators
- From Stateless to Situated: Building a Psychological World for LLM-Based Agents
- Accelerating A/B-Tests with Counterfactual Estimation: Reducing Variance through Policy Overlap
- Data and trained models for "Empirical Evidence of Large Language Model's Influence on Human Spoken Communication"
- Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
- HyMobileAgent: Data-Environment Co-Scaling for Efficient GUI Agents
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
- Discrete Action Space as a Prerequisite for GRPO Convergence in Small-Model Continuous Control
- Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
- Align AI to Dynamic Human-AI Workflows
- Reliability-Aware LLM Alignment from Inconsistent Human Feedback
- Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models
- One Student, Many Teachers: Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context
- On the Limits of Support-Preserving Alignment and Bounded Filtering
- CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents
- TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
- Normalized Rewards for Preference Optimization
- Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
- From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
- Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework
- Large Language Models Hack Rewards, and Society
- Geometry-Guided Constraint Learning for LLM Safety Classification
- Pretraining Language Models on Historical Text
- Robust Critics: Defending LLMs Against Multi-Turn Attacks
- FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation
- Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
- Preference Tuning as Spectral Update Reorganization
- Hallucinations Undermine Trust; Metacognition is a Way Forward
- S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
- Response drift across frontier large language models
- RE-AD: Real-Time Requirement Adherence for Data Labeling
- The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs
- Machine understanding
- Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs
- Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor
- Deterministic Replay for AI Agent Systems
- PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
- Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
- Brainrot: Deskilling and Addiction are Overlooked AI Risks
- "AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs
- Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
- How Open Must Language Models be to Enable Reliable Scientific Inference?
- Some Large Language Models Exhibit Consistent Risk Attitudes
- Removing Sandbagging in LLMs by Training with Weak Supervision
- Why AI systems don't learn and what to do about it: Lessons on autonomous learning from cognitive science
- AI Can Learn Scientific Taste
- Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation
- Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights
- Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
- Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
- Discovering Differences in Strategic Behavior Between Humans and LLMs
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
- The Illusion of Insight in Reasoning Models
- Self-Distillation Enables Continual Learning
- How Human is AI? Examining the Impact of Emotional Prompts on Artificial and Human and Responsiveness
- Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models
- Legal Alignment for Safe and Ethical AI
- Extracting books from production language models
- No Free Lunch in Language Model Bias Mitigation? Targeted Bias Reduction Can Exacerbate Unmitigated LLM Biases
- Distributional AGI Safety
- Reasoning Models Ace the CFA Exams
- Epistemological Fault Lines Between Human and Artificial Intelligence
- Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025
- Echoing: Identity Failures when LLM Agents Talk to Each Other
- Are Large Language Models Sensitive to the Motives Behind Communication?
- SAM 3D: 3Dfy Anything in Images
- Kimi Linear: An Expressive, Efficient Attention Architecture
- A Pragmatic View of AI Personhood
- Everyone prefers human writers, including AI
- LLaDA-MoE: A Sparse MoE Diffusion Language Model
- Video models are zero-shot learners and reasoners
- Layer-0 Suppressors Ground Hallucination Inevitability: A Mechanistic Account of How Transformers Trade Factuality for Hedging
- Understanding Reinforcement Learning for Model Training, and future directions with GRAPE
- An Economy of AI Agents
- K2-Think: A Parameter-Efficient Reasoning System
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design
- On the Theoretical Limitations of Embedding-Based Retrieval
- Measuring Scalar Constructs in Social Science with LLMs
- Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
- Large Language Models Do Not Simulate Human Psychology
- Training language models to be warm and empathetic makes them less reliable and more sycophantic
- Technological folie à deux: Feedback Loops Between AI Chatbots and Mental Illness
- Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
- Relative Entropy Pathwise Policy Optimization
- Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
- LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
- Cognitive models can reveal interpretable value trade-offs in language models
- Who Does What in Deep Learning? Multidimensional Game-Theoretic Attribution of Function of Neural Units
- LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
- Mercury: Ultra-Fast Language Models Based on Diffusion
- InfoFlood: Jailbreaking Large Language Models with Information Overload
- Self-Adapting Language Models
- Unsupervised Elicitation of Language Models
- Reinforcement Pre-Training
- Securing AI Agents with Information-Flow Control
- A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
- Breaking the Performance Ceiling in Reinforcement Learning requires Inference Strategies
- Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods
- RLSR: Reinforcement Learning from Self Reward
- Sycophantic AI decreases prosocial intentions and promotes dependence
- Absolute Zero: Reinforced Self-play Reasoning with Zero Data
- El Agente: An autonomous agent for quantum chemistry
- Do Language Models Know Who Did What to Whom?
- Coral Protocol: Open Infrastructure Connecting The Internet of Agents
- Base Models Beat Aligned Models at Randomness and Creativity
- Kongzi: A Historical Large Language Model with Fact Enhancement
- AssistanceZero: Scalably Solving Assistance Games
- Sugar-Coated Poison: Benign Generation Unlocks LLM Jailbreaking
- VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
- Not All Data Are Unlearned Equally
- Inference-Time Scaling for Generalist Reward Modeling
- A matter of principle? AI alignment as the fair treatment of claims
- Training large language models on narrow tasks can lead to broad misalignment
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
- Large Language Diffusion Models
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
- Rejected Dialects: Biases Against African American Language in Reward Models
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Re-evaluating Theory of Mind evaluation in large language models
- Why human-AI relationships need socioaffective alignment
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- Rethinking Early Stopping: Refine, Then Calibrate
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- A Toolbox for Improving Evolutionary Prompt Search
- The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
- From Accuracy to Impact: The Impact-Driven AI Framework (IDAIF) for Aligning Engineering Architecture with Theory of Change
- The Future of NLP may not be at NLP Conferences: Scholarly Migration Patterns in Natural Language Processing
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories
- Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
- Alpha-R1: Alpha Screening with LLM Reasoning via Reinforcement Learning
- Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
- Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
- C2PO: Diagnosing and Disentangling Bias Shortcuts in LLMs
- Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
- MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
- CORE: A Unified Cascaded Ordinal Relevance Estimation Framework for E-commerce Search
- Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation
- EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion
- REVEALER: Reinforcement-Guided Visual Reasoning for Element-Level Text-Image Alignment Evaluation
- Computational structuralism: Toward a formal theory of meaning in the age of digital intelligence
- Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
- The Reward Model Selection Crisis in Personalized Alignment
- Enhancing knowledge graph interactions: A comprehensive Text-to-Cypher pipeline with large language models
- ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
- Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
- APO: Alpha-Divergence Preference Optimization
- Position: Don't Just "Fix it in Post": A Science of AI Must Study Training Dynamics
- Solipsistic Superintelligence is Unlikely to be Cooperative
- LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
- Eliciting Behaviors in Multi-Turn Conversations
- Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
- ReDiF: Reinforced Distillation for Few Step Diffusion
- Harnessing Large Language Models for Biomedical Named Entity Recognition
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- DICE: Discrete Interpretable Comparative Evaluation with Probabilistic Scoring for Retrieval-Augmented Generation
- M2G-Eval: Enhancing and Evaluating Multi-granularity Multilingual Code Generation
- Lessons from Neuroscience for AI: How integrating Actions, Compositional Structure and Episodic Memory could enable Safe, Interpretable and Human-Like AI
- Learning When Not to Attend Globally
- Role-Based Fault Tolerance System for LLM RL Post-Training
- DarkPatterns-LLM: A Multi-Layer Benchmark for Detecting Manipulative and Harmful AI Behavior
- Pointer-Augmented Autoregressive Generation of Patent Claims with Joint Topology and Content Decoding
- HalluMat: Detecting Hallucinations in LLM-Generated Materials Science Content Through Multi-Stage Verification
- The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
- Learning from Negative Examples: Why Warning-Framed Training Data Teaches What It Warns Against
- Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
- Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
- Unbiased Visual Reasoning with Controlled Visual Inputs
- SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
- Expert-Grounded Automatic Prompt Engineering for Extracting Lattice Constants of High-Entropy Alloys from Scientific Publications using Large Language Models
- ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning
- Understanding Tone-Dependent Inference Cost in Large Language Models
- Intelligence Without Integrity: Why Capable LLMs May Undermine Reliability
- Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning
- LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction
- From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
- Scale Weight Decay and Train Better
- Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug Test
- CONSISTRE: A Unified Consistency-Aware Framework for Document-Level Relation Extraction with Large Language Models
- What do Reward Models Memorize?
- Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes
- Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning
- Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models
- Online Policy Evaluation for MDPs with Dynamic UBSR Measures
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- Moral Hazard in Multi-Agent Language Models
- Security and Privacy in Agentic AI: Grand Challenges and Future Directions
- Reward Guided Decoding for Generative Recommendation
- Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
- Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
- Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
- Inverse RL Helps Align AI by Imitating Humans
- Learning from 53.6K Real-World Developer Edits of AI-Generated Code
- Mask2Shield: Strengthening LLM Safety against Neuron-Pruning Attacks
- In-Context Learning as Implicit Policy Gradient
- Traceable LLM Reasoning for Fake-Order Fraud Detection
- Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining
- Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias
- AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning
- Visual Token Compression Enhances Robustness of MLLMs
- MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation
- Retrieval-based and Fine-tuned LLM Approaches for Industrial Asset Health Monitoring and Decision Support
- How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
- PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF
- STAIF: A Stage-wise Optimization for Complex Instruction Following
- ARdena: Scenario-driven control of real-time LLM agents
- DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training
- Group Preference Collapse in Personalized Multimodal Large Language Models
- RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection
- Codifying the Judge: Scalable Evaluation via Program Distillation
- Latent Space Probing for Adult Content Detection in Video Generative Models
- Context Sensitivity Improves Human-Machine Visual Alignment
- The Cartesian Cut in Agentic AI
- Can an Actor-Critic Optimization Framework Improve Analog Design?
- LanteRn: Latent Visual Structured Reasoning
- SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
- Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
- Responsible Intelligence in Practice: A Fairness Audit of Open Large Language Models for Library Reference Services
- El Agente Estructural: An Artificially Intelligent Molecular Editor
- Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis
- SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams
- Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models
- DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
- A Comedy of Estimators: On KL Regularization in RL Training of LLMs
- Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures
- GoldenFuzz: Generative Golden Reference Hardware Fuzzing
- RLLaVA: An RL-central Framework for Language and Vision Assistants
- Beyond Context: Large Language Models' Failure to Grasp Users' Intent
- Semi-Supervised Learning for Large Language Models Safety and Content Moderation
- Artificial or Just Artful? Do LLMs Bend the Rules in Programming?
- GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs
- DIAL: Direct Iterative Adversarial Learning for Realistic Multi-Turn Dialogue Simulation
- Generalization of RLVR Using Causal Reasoning as a Testbed
- FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
- Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
- BRIDGE: Budget-aware Reasoning via Intermediate Distillation with Guided Examples
- Offline Safe Policy Optimization From Heterogeneous Feedback
- Learning to Reason in LLMs by Expectation Maximization
- Memory-T1: Reinforcement Learning for Temporal Reasoning in Multi-session Agents
- Counterfactual LLM-based Framework for Measuring Rhetorical Style
- Scalable Stewardship of an LLM-Assisted Clinical Benchmark with Physician Oversight
- Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation
- Learning General Policies with Policy Gradient Methods
- Humanlike AI Design Increases Anthropomorphism but Yields Divergent Outcomes on Engagement and Trust Globally
- QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation
- AWPO: Enhancing Tool-Use of Large Language Models through Adaptive Integration of Reasoning Rewards
- Explicit and Non-asymptotic Query Complexities of Rank-Based Zeroth-order Algorithm on Stochastic Smooth Functions
- Population-Evolve: a Parallel Sampling and Evolutionary Method for LLM Math Reasoning
- Efficient Personalization of Generative Models via Optimal Experimental Design
- Recontextualization Mitigates Specification Gaming without Modifying the Specification
- ORPR: An OR-Guided Pretrain-then-Reinforce Learning Model for Inventory Management
- Online Robust Reinforcement Learning with General Function Approximation
- FASTRIC: Prompt Specification Language for Verifiable LLM Interactions
- Auto-Prompting with Retrieval Guidance for Frame Detection in Logistics
- MEEA: Mere Exposure Effect-Driven Confrontational Optimization for LLM Jailbreaking
- SafeMed-R1: Adversarial Reinforcement Learning for Generalizable and Robust Medical Reasoning in Vision-Language Models
- Learning Hierarchical Procedural Memory for LLM Agents through Bayesian Selection and Contrastive Refinement
- RMLer: Synthesizing Novel Objects across Diverse Categories via Reinforcement Mixing Learning
- HARBOR: Holistic Adaptive Risk assessment model for BehaviORal healthcare
- Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
- LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction
- LLMs on Drugs: Language Models Are Few-Shot Consumers
- A Multi-agent Text2SQL Framework using Small Language Models and Execution Feedback
- Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
- Software Vulnerability Management in the Era of Artificial Intelligence: An Industry Perspective
- CrystalFormer-CSP: Thinking Fast and Slow for Crystal Structure Prediction
- Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation
- Stable and Efficient Single-Rollout RL for Multimodal Reasoning
- Trustworthy and Explainable Deep Reinforcement Learning for Safe and Energy-Efficient Process Control: A Use Case in Industrial Compressed Air Systems
- ReGal: A First Look at PPO-based Legal AI for Judgment Prediction and Summarization in India
- Adversarial Robustness of Vision in Open Foundation Models
- ShareChat: A Dataset of Chatbot Conversations in the Wild
- AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
- Subjective Question Generation and Answer Evaluation using NLP
- Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
- Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL Finetuning
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
- Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topics
- Synthelite: Chemist-aligned and feasibility-aware synthesis planning with LLMs
- Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
- Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
- MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
- PAACE: A Plan-Aware Automated Agent Context Engineering Framework
- Small Language Models for Efficient Agentic Tool Calling: Outperforming Large Models with Targeted Fine-tuning
- Towards Proactive Personalization through Profile Customization for Individual Users in Dialogues
- Stepwise Think-Critique: A Unified Framework for Robust and Interpretable LLM Reasoning
- Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
- Evaluating LLMs for Zeolite Synthesis Event Extraction (ZSEE): A Systematic Analysis of Prompting Strategies
- MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
- The Semantic Illusion: Certified Limits of Embedding-Based Hallucination Detection in RAG Systems
- Spectral Representation-based Reinforcement Learning
- Model Agnostic Preference Optimization for Medical Image Segmentation
- DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
- Can LLMs Guide Their Own Exploration? Gradient-Guided Reinforcement Learning for LLM Reasoning
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections
- Entropy-Reservoir Bregman Projection: An Information-Geometric Unification of Model Collapse
- Learning to Extract Context for Context-Aware LLM Inference
- IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
- Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
- Inflation Attitudes of Large Language Models
- Georeferencing complex relative locality descriptions with large language models
- History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation
- Explainable Ethical Assessment on Human Behaviors by Generating Conflicting Social Norms
- Understanding and Improving Hyperbolic Deep Reinforcement Learning
- Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
- A First-Order Logic-Based Alternative to Reward Models in RLHF
- A Multifaceted Analysis of Social Biases in Large Language Models
- OpenDataArena: A Fair and Open Arena for Benchmarking Post-Training Dataset Value
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
- Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation
- Explainable reinforcement learning from human feedback to improve alignment
- CAPE: Capability Achievement via Policy Execution
- Towards Effective Model Editing for LLM Personalization
- Towards Interactive Intelligence for Digital Humans
- A Scientific Reasoning Model for Organic Synthesis Procedure Generation
- State-Dependent Refusal and Learned Incapacity in RLHF-Aligned Language Models
- MedCEG: Reinforcing Verifiable Medical Reasoning with Critical Evidence Graph
- MiniLingua: A Small Open-Source LLM for European Languages
- Differentiable Evolutionary Reinforcement Learning
- AIR: Post-training Data Selection for Reasoning via Attention Head Influence
- Post-Training and Test-Time Scaling of Generative Agent Behavior Models for Interactive Autonomous Driving
- Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
- Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
- SpeakRL: Synergizing Reasoning, Speaking, and Acting in Language Models with Reinforcement Learning
- Socratic Students: Teaching Language Models to Learn by Asking Questions
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
- Revisiting the Reliability of Language Models in Instruction-Following
- What Matters in Evaluating Book-Length Stories? A Systematic Study of Long Story Evaluation
- Anchoring Values in Temporal and Group Dimensions for Flow Matching Model Alignment
- How Prompts Move Language Model Behavior: Frames, Salience, and Construal as Semantic Control
- HyperEdit: Unlocking Instruction-based Text Editing in LLMs via Hypernetworks
- Adaptive Detector-Verifier Framework for Zero-Shot Polyp Detection in Open-World Settings
- Unified Control for Inference-Time Guidance of Denoising Diffusion Models
- The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior
- Taxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Models
- Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
- Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization
- Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
- Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models
- RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
- ReactorFold: Generative discovery of nuclear reactor cores via emergent physical reasoning
- Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
- Causal Judge Evaluation: Calibrated Surrogate Metrics for LLM Systems
- MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents
- Your plan may succeed, but what about failure? Investigating how people use ChatGPT for long-term life task planning
- TAO-Net: Two-stage Adaptive OOD Classification Network for Fine-grained Encrypted Traffic Classification
- PIAST: Rapid Prompting with In-context Augmentation for Scarce Training data
- RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems
- LLM-Auction: Generative Auction towards LLM-Native Advertising
- Multi-dimensional Preference Alignment by Conditioning Reward Itself
- Exploring LLMs for Scientific Information Extraction Using The SciEx Framework
- A Unified Generative-Predictive Framework for Deterministic Inverse Design
- KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question Answering
- MedBioRAG: Semantic Search and Retrieval-Augmented Generation with Large Language Models for Medical and Biological QA
- Rethinking Chain-of-Thought Reasoning for Videos
- System Report for CCL25-Eval Task 10: Prompt-Driven Large Language Model Merge for Fine-Grained Chinese Hate Speech Detection
- Building Reasonable Inference for Vision-Language Models in Blind Image Quality Assessment
- Chasing Shadows: Pitfalls in LLM Security Research
- RouteRAG: Efficient Retrieval-Augmented Generation from Text and Graph via Reinforcement Learning
- Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment
- Encoder-Free Knowledge-Graph Reasoning with LLMs via Hyperdimensional Path Retrieval
- A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
- Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
- Gradient-Informed Monte Carlo Fine-Tuning of Diffusion Models for Low-Thrust Trajectory Design
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- Uncertainty-Aware Data-Efficient AI: An Information-Theoretic Perspective
- rSIM: Incentivizing Reasoning Capabilities of LLMs via Reinforced Strategy Injection
- A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties
- Universal Adversarial Suffixes for Language Models Using Reinforcement Learning with Calibrated Reward
- Large Language Models for Education and Research: An Empirical and User Survey-based Analysis
- Provable Long-Range Benefits of Next-Token Prediction
- Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
- SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
- Depth-Wise Activation Steering for Honest Language Models
- MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
- ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
- ContextualSHAP : Enhancing SHAP Explanations Through Contextual Language Generation
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
- Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
- Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels
- VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection
- The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
- Parent-Guided Semantic Reward Model (PGSRM): Embedding-Based Reward Functions for Reinforcement Learning of Transformer Language Models
- Rhea: Role-aware Heuristic Episodic Attention for Conversational LLMs
- Agency at the Interface: Distinguishing Teleological from Structural Self-Organization via Internal Coarse-Graining and Downward Causation
- SIEVE: Selective Integrity Verification and Escalation for Defending LLM Agents against Indirect Prompt Injection
- Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generation
- Personalized Image Descriptions from Attention Sequences
- PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
- RLAX: Large-Scale, Distributed Reinforcement Learning for Large Language Models on TPUs
- When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
- Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models
- Auto-exploration for online reinforcement learning
- Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge
- Beyond Token-level Supervision: Unlocking the Potential of Decoding-based Regression via Reinforcement Learning
- ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language Models
- A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
- ProSocialAlign: Preference Conditioned Test Time Alignment in Language Models
- ARCANE: A Multi-Agent Framework for Interpretable and Configurable Alignment
- PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
- LLM Harms: A Taxonomy and Discussion
- Intrinsically Interpretable Attention via Sparse Post-Training
- AGI Requires a Coordination Layer on Top of Pattern Repositories
- ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
- Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning
- The Road of Adaptive AI for Precision in Cybersecurity
- SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
- Dynamic Alignment for Collective Agency: Toward a Scalable Self-Improving Framework for Open-Ended LLM Alignment
- Value Gradient Guidance for Flow Matching Alignment
- LMSpell: Neural Spell Checking for Low-Resource Languages
- AI & Human Co-Improvement for Safer Co-Superintelligence
- STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models
- STELLA: Guiding Large Language Models for Time Series Forecasting with Semantic Abstractions
- SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge Graphs
- Reflection-Satisfaction Tradeoff: Investigating Impact of Reflection on Student Engagement with AI-Generated Programming Hints
- YingMusic-Singer: Zero-shot Singing Voice Synthesis and Editing with Annotation-free Melody Guidance
- RLHFSpec: Breaking the Efficiency Bottleneck in RLHF Training via Adaptive Drafting
- Towards an AI Fluid Scientist: LLM-Powered Scientific Discovery in Experimental Fluid Mechanics
- Efficient Reinforcement Learning with Semantic and Token Entropy for LLM Reasoning
- Cross-Task Benchmarking and Evaluation of General-Purpose and Code-Specific Large Language Models
- Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective
- ADAPT: Learning Task Mixtures for Budget-Constrained Instruction Tuning
- VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
- ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
- Data-regularized Reinforcement Learning for Diffusion Models at Scale
- Towards better dense rewards in Reinforcement Learning Applications
- Bootstrapped Mixed Rewards for RL Post-Training: Injecting Canonical Action Order
- Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
- Enhancing Instruction-Following Capabilities in Seq2Seq Models: DoLA Adaptations for T5
- In-Context Representation Hijacking
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- FFTrainer: Fast Failover in Large-Language Model Training with Almost-Free State Management
- Us-vs-Them bias in Large Language Models
- Don't Trust Your Upstream: Exploiting LLM Multi-Agent System via Topology-Guided Adversarial Propagation
- Full-Stack Alignment: Co-Aligning AI and Institutions with Thick Models of Value
- From static to adaptive: immune memory-based jailbreak detection for large language models
- Idea-Gated Transformers: Enforcing Semantic Coherence via Differentiable Vocabulary Pruning
- PretrainZero: Reinforcement Active Pretraining
- MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking
- Invasive Context Engineering to Control Large Language Models
- Hypothesis Testing for Generalized Thurstone Models
- Network Self-Configuration based on Fine-Tuned Small Language Models
- SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment
- Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models
- Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics Shifts
- Guided Self-Evolving LLMs with Minimal Human Supervision
- OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
- promptolution: A Unified, Modular Framework for Prompt Optimization
- From monoliths to modules: Decomposing transducers for efficient world modelling
- Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning
- AlignSAE: Concept-Aligned Sparse Autoencoders
- Artemis: Structured Visual Reasoning for Perception Policy Learning
- Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
- GrndCtrl: Grounding World Models via Self-Supervised Reward Alignment
- Agentic Policy Optimization via Instruction-Policy Co-Evolution
- Rectifying LLM Thought from Lens of Optimization
- Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
- Exploring Human Perceptions of AI Responses: Insights from a Mixed-Methods Study on Risk Mitigation in Generative Models
- OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
- CauSight: Learning to Supersense for Visual Causal Discovery
- How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
- Zero-Overhead Introspection for Adaptive Test-Time Compute
- AI-Enabled grading with near-domain data for scaling feedback with human-level accuracy
- PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards
- Financial Instruction Following Evaluation (FIFE)
- S2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
- When Safety Blocks Sense: Measuring Semantic Confusion in LLM Refusals
- On The Finetuning of MLIPs Through the Lens of Iterated Maps With BPTT
- Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
- Towards Active Synthetic Data Generation for Finetuning Language Models
- ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
- Clinical-R1: Empowering Large Language Models for Faithful and Comprehensive Reasoning with Clinical Objective Relative Policy Optimization
- Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
- UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
- SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
- When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
- Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation
- Aligning Probabilistic Beliefs under Informative Missingness: LLM Steerability in Clinical Reasoning
- EduEval: A Hierarchical Cognitive Benchmark for Evaluating Large Language Models in Chinese Education
- Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
- Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization
- Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards
- Instruction Tuning of Large Language Models for Tabular Data Generation-in One Day
- Listwise Preference Optimization with Element-wise Confusions for Aspect Sentiment Quad Prediction
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
- OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving
- Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
- Co-Evolving Agents: Learning from Failures as Hard Negatives
- Optimizing NetGPT via Routing-Based Synergy and Reinforcement Learning
- Tacit Bidder-Side Collusion: Artificial Intelligence in Dynamic Auctions
- TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
- PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and Fuzz Optimization
- Decomposed Trust: Exploring Privacy, Adversarial Robustness, Fairness, and Ethics of Low-Rank LLMs
- MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
- Video Generation Models Are Good Latent Reward Models
- Optimizing Life Sciences Agents in Real-Time using Reinforcement Learning
- Bootstrapping LLMs via Preference-Based Policy Optimization
- The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training
- GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs
- Can Finetuing LLMs on Small Human Samples Increase Heterogeneity, Alignment, and Belief-Action Coherence?
- Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models
- TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models
- OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
- ICPO: Intrinsic Confidence-Driven Group Relative Preference Optimization for Efficient Reinforcement Learning
- Escaping the Verifier: Learning to Reason via Demonstrations
- Towards Audio Token Compression in Large Audio Language Models
- Reinforcement Learning for Latent-Space Thinking in LLMs
- Reinforcing Action Policies by Prophesying
- MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models
- On Evaluating LLM Alignment by Evaluating LLMs as Judges
- DesignPref: Capturing Personal Preferences in Visual Design Generation
- DRAFT-RL: Multi-Agent Chain-of-Draft Reasoning for Reinforcement Learning-Enhanced LLMs
- Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
- Thinking in 360°: Humanoid Visual Search in the Wild
- AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models
- CREward: A Type-Specific Creativity Reward Model
- HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
- Differential Smoothing Mitigates Sharpening and Improves LLM Reasoning
- Profile-LLM: Dynamic Profile Optimization for Realistic Personality Expression in LLMs
- CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
- Learning to Clean: Reinforcement Learning for Noisy Label Correction
- SOMBRL: Scalable and Optimistic Model-Based RL
- SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space
- Advances and Challenges in Solar Flare Prediction: A Review
- Schema Matching on Graph: Iterative Graph Exploration for Efficient and Explainable Data Integration
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
- EAGER: Edge-Aligned LLM Defense for Robust, Efficient, and Accurate Cybersecurity Question Answering
- Large Language Models as Search Engines: Societal Challenges
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Beyond Reward Margin: Rethinking and Resolving Likelihood Displacement in Diffusion Models via Video Generation
- OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
- Medical Malice: A Dataset for Context-Aware Safety in Healthcare LLMs
- Accelerating Reinforcement Learning via Error-Related Human Brain Signals
- Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
- FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
- Optimizing LLM Code Suggestions: Feedback-Driven Timing with Lightweight State Bounds
- RoguePrompt: Dual-Layer Ciphering for Self-Reconstruction to Circumvent LLM Moderation
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation
- Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation
- Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
- Test-Time Preference Optimization for Image Restoration
- Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
- TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization
- Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI
- Exploring Weak-to-Strong Generalization for CLIP-based Classification
- Building Domain-Specific Small Language Models via Guided Data Generation
- Efficient Inference Using Large Language Models with Limited Human Data: Fine-Tuning then Rectification
- Curvature-Aware Safety Restoration In LLMs Fine-Tuning
- SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
- Alignment Faking - the Train -> Deploy Asymmetry: Through a Game-Theoretic Lens with Bayesian-Stackelberg Equilibria
- MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
- Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
- QuickLAP: Quick Language-Action Preference Learning for Semi-Autonomous Agents
- Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently
- MultiGA: Leveraging Multi-Source Seeding in Genetic Algorithms
- Mesh RAG: Retrieval Augmentation for Autoregressive Mesh Generation
- The PLLuM Instruction Corpus
- PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
- Why Do Language Model Agents Whistleblow?
- Supervised Fine Tuning of Large Language Models for Domain Specific Knowledge Graph Construction:A Case Study on Hunan's Historical Celebrities
- FIRM: Federated In-client Regularized Multi-objective Alignment for Large Language Models
- Evaluating Adversarial Vulnerabilities in Modern Large Language Models
- Neighbor GRPO: Contrastive ODE Policy Optimization Aligns Flow Models
- Cross-cultural value alignment frameworks for responsible AI governance: Evidence from China-West comparative analysis
- MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
- SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation
- Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
- Personalized Reward Modeling for Text-to-Image Generation
- Towards Unified Vision Language Models for Forest Ecological Analysis in Earth Observation
- SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning
- "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
- Pass@k Metric for RLVR: A Diagnostic Tool of Exploration, But Not an Objective
- PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
- Multi-Agent Collaborative Reward Design for Enhancing Reasoning in Reinforcement Learning
- Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models
- A Mathematical Framework for Custom Reward Functions in Job Application Evaluation using Reinforcement Learning
- Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning
- The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment
- Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Reflexive Evidence-Based Multimodal Learning for Clean Energy Transitions: Causal Insights on Cooking Fuel Access, Urbanization, and Carbon Emissions
- Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learning
- Efficiency Will Not Lead to Sustainable Reasoning AI
- GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning
- EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
- SafeRBench: Dissecting the Reasoning Safety of Large Language Models
- Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
- Learning Human-Like RL Agents Through Trajectory Optimization With Action Quantization
- GeoSceneGraph: Geometric Scene Graph Diffusion Model for Text-guided 3D Indoor Scene Synthesis
- Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
- GPS: General Per-Sample Prompter
- Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior
- Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning
- Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distribution
- Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
- Just Asking Questions: Doing Our Own Research on Conspiratorial Ideation by Generative AI Chatbots
- Retrieval-GRPO: A Multi-Objective Reinforcement Learning Framework for Dense Retrieval in Taobao Search
- VLMs Guided Interpretable Decision Making for Autonomous Driving
- Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
- Generalist Foundation Models Are Not Clinical Enough for Hospital Operations
- ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models
- Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
- Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance
- Evaluating the Ability of Large Language Models to Identify Adherence to CONSORT Reporting Guidelines in Randomized Controlled Trials: A Methodological Evaluation Study
- BeDiscovER: The Benchmark of Discourse Understanding in the Era of Reasoning Language Models
- STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
- Learning from the Undesirable: Robust Adaptation of Language Models without Forgetting
- The Future of Food: How Artificial Intelligence is Transforming Food Manufacturing
- RLHF May Not Reflect Genuine Preferences
- The Alignment Game: A Theory of Long-Horizon Alignment Through Recursive Curation
- LLM Reinforcement in Context
- Reg-DPO: SFT-Regularized Direct Preference Optimization with GT-Pair for Improving Video Generation
- Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
- RAGSmith: A Framework for Finding the Optimal Composition of Retrieval-Augmented Generation Methods Across Datasets
- Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing
- Group-Aware Reinforcement Learning for Output Diversity in Large Language Models
- Detecting LLM-Assisted Academic Dishonesty using Keystroke Dynamics
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
- Tailored Primitive Initialization is the Secret Key to Reinforcement Learning
- From Passive to Persuasive: Localized Activation Injection for Empathy and Negotiation
- Maximizing the efficiency of human feedback in AI alignment: a comparative analysis
- Mitigating Length Bias in RLHF through a Causal Lens
- AlignTree: Efficient Defense Against LLM Jailbreak Attacks
- Multi-Value Alignment for LLMs via Value Decorrelation and Extrapolation
- Rethinking Deep Alignment Through The Lens Of Incomplete Learning
- MetaGDPO: Alleviating Catastrophic Forgetting with Metacognitive Knowledge through Group Direct Preference Optimization
- EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generation
- On the Entropy Calibration of Language Models
- PIRA: Preference-Oriented Instruction-Tuned Reward Models with Dual Aggregation
- A Multifaceted Analysis of Negative Bias in Large Language Models through the Lens of Parametric Knowledge
- Context-Emotion Aware Therapeutic Dialogue Generation: A Multi-component Reinforcement Learning Approach to Language Models for Mental Health Support
- Aligning Machiavellian Agents: Behavior Steering via Test-Time Policy Shaping
- Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
- Bytes of a Feather: Personality and Opinion Alignment Effects in Human-AI Interaction
- EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
- VIDEOP2R: Video Understanding from Perception to Reasoning
- Dynamic Temperature Scheduler for Knowledge Distillation
- From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
- Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
- AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
- Persona-Aware Alignment Framework for Personalized Dialogue Generation
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models
- In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback
- Uncertainty-Guided Checkpoint Selection for Reinforcement Finetuning of Large Language Models
- Black-Box On-Policy Distillation of Large Language Models
- Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
- LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
- SlideBot: A Multi-Agent Framework for Generating Informative, Reliable, Multi-Modal Presentations
- NSL-MT: Linguistically Informed Negative Samples for Efficient Machine Translation in Low-Resource Languages
- Where does an LLM begin computing an instruction?
- AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
- C3TG: Conflict-aware, Composite, and Collaborative Controlled Text Generation
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
- LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
- RGMP: Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation
- Thinking Forward and Backward: Multi-Objective Reinforcement Learning for Retrieval-Augmented Reasoning
- Automatic Minds: Cognitive Parallels Between Hypnotic States and Large Language Model Processing
- Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
- HalluClean: A Unified Framework to Combat Hallucinations in LLMs
- Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
- The Path Not Taken: RLVR Provably Learns Off the Principals
- Reinforcement Learning Control of Quantum Error Correction
- How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
- Patching LLM Like Software: A Lightweight Method for Improving Safety Policy in Large Language Models
- DPRM: A Dual Implicit Process Reward Model in Multi-Hop Question Answering
- AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
- Prompt Tuning for Natural Language to SQL with Embedding Fine-Tuning and RAG
- Safety-Preserving PTQ via Contrastive Alignment Loss
- DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
- SERL: Self-Examining Reinforcement Learning on Open-Domain
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
- Majority Rules: LLM Ensemble is a Winning Approach for Content Categorization
- Distributionally Robust Online Markov Game with Linear Function Approximation
- PC-Diffusion: Aligning Diffusion Models with Human Preferences via Preference Classifier
- Judging by the Rules: Compliance-Aligned Framework for Modern Slavery Statement Monitoring
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
- Cortex AISQL: A Production SQL Engine for Unstructured Data
- A Self-Improving Architecture for Dynamic Safety in Large Language Models
- On the Creativity of AI Agents
- What can LLMs tell us about the mechanisms behind polarity illusions in humans? Experiments across model scales and training steps
- RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services
- Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
- HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection
- EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
- Learning to Focus: Focal Attention for Selective and Scalable Transformers
- Textual Self-attention Network: Test-Time Preference Optimization through Textual Gradient-based Attention
- SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
- MENTOR: A Metacognition-Driven Self-Evolution Framework for Uncovering and Mitigating Implicit Domain Risks in LLMs
- RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
- FedRW: Efficient Privacy-Preserving Data Reweighting for Enhancing Federated Learning of Language Models
- Convergence of Actor-Critic Learning for Mean Field Games and Mean Field Control in Continuous Spaces
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
- EASE: Practical and Efficient Safety Alignment for Small Language Models
- What Makes Reasoning Invalid: Echo Reflection Mitigation for Large Language Models
- Adaptive Regularization for Large-Scale Sparse Feature Embedding Models
- Synthetic Data-Driven Prompt Tuning for Financial QA over Tables and Documents
- OpenVLN: Open-world Aerial Vision-Language Navigation
- KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
- Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data
- Kunlun Anomaly Troubleshooter: Enabling Kernel-Level Anomaly Detection and Causal Reasoning for Large Model Distributed Inference
- L2T-Hyena: Enhancing State-Space Models with an Adaptive Learn-to-Teach Framework
- Lived Experience in Dialogue: Co-designing Personalization in Large Language Models to Support Youth Mental Well-being
- A Representation Sharpening Framework for Zero Shot Dense Retrieval
- RLoop: An Self-Improving Framework for Reinforcement Learning with Iterative Policy Initialization
- Optimizing Diversity and Quality through Base-Aligned Model Collaboration
- Steering Language Models with Weight Arithmetic
- PreResQ-R1: Towards Fine-Grained Rank-and-Score Reinforcement Learning for Visual Quality Assessment via Preference-Response Disentangled Policy Optimization
- Reasoning on Time-Series for Financial Technical Analysis
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- Building Specialized Software-Assistant ChatBot with Graph-Based Retrieval-Augmented Generation
- Reflective Personalization Optimization: A Post-hoc Rewriting Framework for Black-Box Large Language Models
- Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
- Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning
- You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models
- Explore Data Left Behind in Reinforcement Learning for Reasoning Language Models
- CPO: Condition Preference Optimization for Controllable Image Generation
- Personalized Image Editing in Text-to-Image Diffusion Models via Collaborative Direct Preference Optimization
- Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
- Black-Box Guardrail Reverse-engineering Attack
- Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs
- Advancing Equitable AI: Evaluating Cultural Expressiveness in LLMs for Latin American Contexts
- MIDI-LLM: Improving Text-to-MIDI Music Generation via Adapting Large Language Models
- RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
- SynQuE: Estimating Synthetic Dataset Quality Without Annotations
- Test-Time Adaptation for LLM Agents via Environment Interaction
- GRAD: Graph-Retrieved Adaptive Decoding for Hallucination Mitigation
- STARS: Synchronous Token Alignment for Robust Supervision in Large Language Models
- Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology
- Learning Without Critics? Revisiting GRPO in Classical Reinforcement Learning Environments
- LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
- DeepKnown-Guard: A Proprietary Model-Based Safety Response Framework for AI Agents
- Control Barrier Function for Aligning Large Language Models
- COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
- Silenced Biases: The Dark Side LLMs Learned to Refuse
- A Collaborative Reasoning Framework for Anomaly Diagnostics in Underwater Robotics
- Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMs
- Epidemiology of Large Language Models: A Benchmark for Observational Distribution Knowledge
- Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions
- PublicAgent: Multi-Agent Design Principles From an LLM-Based Open Data Analysis Framework
- AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence
- Controlling Performance and Budget of a Centralized Multi-agent LLM System with Reinforcement Learning
- Extending RLVR to Open-Ended Tasks via Verifiable Multiple-Choice Reformulation
- Understanding New-Knowledge-Induced Factual Hallucinations in LLMs: Analysis and Interpretation
- The Realignment Problem: When Right becomes Wrong in LLMs
- Directional-Clamp PPO
- AyurParam: A State-of-the-Art Bilingual Language Model for Ayurveda
- Can Conversational AI Counsel for Change? A Theory-Driven Approach to Supporting Dietary Intentions in Ambivalent Individuals
- An Automated Framework for Strategy Discovery, Retrieval, and Evolution in LLM Jailbreak Attacks
- SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
- LLMs as Judges: Toward The Automatic Review of GSN-compliant Assurance Cases
- Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
- Inference-Time Personalized Alignment with a Few User Preference Queries
- Automated Reward Design for Gran Turismo
- DL4Proteins Jupyter Notebooks Teach how to use Artificial Intelligence for Biomolecular Structure Prediction and Design
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow Preferences
- Random Initialization of Gated Sparse Adapters
- RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
- 3EED: Ground Everything Everywhere in 3D
- Efficient Test-Time Retrieval Augmented Generation
- DPO-F+: Aligning Code Repair Feedback with Developers' Preferences
- Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learning
- IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation
- Do Math Reasoning LLMs Help Predict the Impact of Public Transit Events?
- A Systematic Literature Review of Code Hallucinations in LLMs: Characterization, Mitigation Methods, Challenges, and Future Directions for Reliable AI
- Ariadne: A Controllable Framework for Probing and Extending VLM Reasoning Boundaries
- DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching
- Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- Reversal Invariance in Autoregressive Language Models
- Reimagining Safety Alignment with An Image
- Diverse Human Value Alignment for Large Language Models via Ethical Reasoning
- VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision
- Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents
- Disrupting Networks: Amplifying Social Dissensus via Opinion Perturbation and Large Language Models
- Characterizing Selective Refusal Bias in Large Language Models
- Addressing Longstanding Challenges in Cognitive Science with Language Models
- BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
- Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
- Closing the Expression Gap in LLM Instructions via Socratic Questioning
- MolChord: Structure-Sequence Alignment for Protein-Guided Drug Design
- Reasoning Up the Instruction Ladder for Controllable Language Models
- Kad: A Framework for Proxy-based Test-time Alignment with Knapsack Approximation Deferral
- FlowMesh: A Service Fabric for Composable LLM Workflows
- SteerVLM: Robust Model Control through Lightweight Activation Steering for Vision Language Models
- Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
- Value Drifts: Tracing Value Alignment During LLM Post-Training
- LoRAQuant: Mixed-Precision Quantization of LoRA to Ultra-Low Bits
- Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
- Data-Efficient RLVR via Off-Policy Influence Guidance
- OmniEduBench: A Comprehensive Chinese Benchmark for Evaluating Large Language Models in Education
- BOTS: A Unified Framework for Bayesian Online Task Selection in LLM Reinforcement Finetuning
- Offline Clustering of Preference Learning with Active-data Augmentation
- Graph-Enhanced Policy Optimization in LLM Agent Training
- Test-Time Alignment of LLMs via Sampling-Based Optimal Control in pre-logit space
- Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
- Similarity-Distance-Magnitude Language Models
- Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
- One Model to Critique Them All: Rewarding Agentic Tool-Use via Efficient Reasoning
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
- LLMBisect: Breaking Barriers in Bug Bisection with A Comparative Analysis Pipeline
- Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
- Approximating Human Preferences Using a Multi-Judge Learned System
- The Information-Theoretic Imperative: Compression and the Epistemic Foundations of Intelligence
- ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
- EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis
- Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves
- Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs
- Not ready for the bench: LLM legal interpretation is unstable and out of step with human judgments
- Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines
- MemSFT: Mitigating Alignment Tax with an External Parametric Memory
- NormWorlds-CF: Solver-Verified Counterfactual Normative Reasoning with Metamorphic-Relation GRPO
- BioDisclose: An Actionability-Aware Benchmark for Biomedical Safety under Adversarial Elicitation
- RecoReward: Recommender-Guided Multimodal Description Generation for Recommendation
- DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
- Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation
- CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
- Deep neural networks and humans both benefit from compositional language structure
- Fairness and Bias in Algorithmic Hiring: A Multidisciplinary Survey
- LLM-Augmented Computational Phenotyping of Long Covid
- Multi-Decoder OneRec: Controllable Generative Retrieval for Multi-Objective Industrial Recommendation
- HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models
- Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
- Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
- Scientific Knowledge Discovery in the Age of Large Language Models
- Constitutional Midtraining: Content Presence Drives Alignment Gains
- Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
- Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
- Embodied Agents Take Control: Minimal-Interface Zero-Shot Agents Rival Industrial-Scale Policies in Vision-and-Language Navigation
- The Innate Economic Preferences of Language Models
- Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback
- SciMON: Scientific Inspiration Machines Optimized for Novelty
- Simulating Subjects: The Promise and Peril of Artificial Intelligence Stand-Ins for Social Agents and Interactions
- RAPID: An Efficient Reinforcement Learning Algorithm for Small Language Models
- Large Language Models for Software Engineering: A Systematic Literature Review
- Large language models propagate race-based medicine
- Enhancing Hate Speech Detection with Fine-Tuned Large Language Models Requires High-Quality Data
- Beware of botshit: How to manage the epistemic risks of generative chatbots
- Aligning Pedagogy with Generative AI: An Approach to Customizing Educational GPTs
- Developing Students’ Statistical Expertise Through Writing in the Age of AI
- Take Goodhart Seriously: Principled Limit on General-Purpose AI Optimization
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- Reward Models are Metrics in a Trench Coat
- Truth-Aware Decoding: A Program-Logic Approach to Factual Language Generation
- TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
- Can Large Language Models Transform Computational Social Science?
- OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment
- Latent-IM: Latent Interaction Management for Speech LLMs
- DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models
- Safety from Honesty in a Disinterested AI Predictor
- The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
- Vision-Language Models Suppress Female Representations Under Ambiguous Input
- ChatGPT and me: First-time and experienced users’ perceptions of ChatGPT’s communicative ability as a dialogue partner
- Learning by teaching with ChatGPT : The effect of teachable ChatGPT agent on programming education
- Large language models for biomedicine: foundations, opportunities, challenges, and best practices
- VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
- Process Matters more than Output for Distinguishing Humans from Machines
- Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning
- How Can Recommender Systems Benefit from Large Language Models: A Survey
- State-Dependent Safety Failures in Multi-Turn Language Model Interaction
- Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
- Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
- Fine-Tuning GPT-5 for GPU Kernel Generation
- Urban-R1: Reinforced MLLMs Mitigate Geospatial Biases for Urban General Intelligence
- LANPO: Bootstrapping Language and Numerical Feedback for Reinforcement Learning in LLMs
- Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
- RL makes MLLMs see better than SFT
- Large language models encode clinical knowledge
- On the Use of Large Language Models for Qualitative Synthesis
- Agents at Risk: How Users Unwittingly Undermine LLM Safety
- Ministral 3
- Investigating the Validity Evidence of Automated Scoring Methods for Divergent Thinking Assessments
- LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs
- Dense and Diverse Goal Coverage in Multi Goal Reinforcement Learning
- PRESTO: Preimage-Informed Instruction Optimization for Prompting Black-Box LLMs
- Sharpness-Guided Group Relative Policy Optimization via Probability Shaping
- Ideology-Based LLMs for Content Moderation
- Agentic Moderation: Multi-Agent Design for Safer Vision-Language Models
- DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA
- Model-Document Protocol for AI Search
- LISTEN to Your Preferences: An LLM Framework for Multi-Objective Selection
- A Survey on Unlearning in Large Language Models
- DEBATE: A Large-Scale Benchmark for Role-Playing LLM Agents in Multi-Agent, Long-Form Debates
- Learning-Based vs Human-Derived Congestion Control: An In-Depth Experimental Study
- Reasoning-Aware GRPO using Process Mining
- FT-ARM: Fine-Tuned Agentic Reflection Multimodal Language Model for Pressure Ulcer Severity Classification with Reasoning
- Greedy Sampling Is Provably Efficient for RLHF
- SPICE: Self-Play In Corpus Environments Improves Reasoning
- OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
- LoRA-DA: Data-Aware Initialization for Low-Rank Adaptation via Asymptotic Analysis
- Critique-RL: Training Language Models for Critiquing through Two-Stage Reinforcement Learning
- Towards Transparent Reasoning: What Drives Faithfulness in Large Language Models?
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- PFEA: An LLM-based High-Level Natural Language Planning and Feedback Embodied Agent for Human-Centered AI
- Semi-Supervised Preference Optimization with Limited Feedback
- MASPRM: Multi-Agent System Process Reward Model
- The Sign Estimator: LLM Alignment in the Face of Choice Heterogeneity
- World Simulation with Video Foundation Models for Physical AI
- Fortytwo: Swarm Inference with Peer-Ranked Consensus
- Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation
- Towards AI as Colleagues: Multi-Agent System Improves Structured Professional Ideation
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- GIFT: Group-relative Implicit Fine Tuning Integrates GRPO with DPO and UNA
- Assessing the Relational Abilities of Large Language Models and Large Reasoning Models
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- On the Impossibility of Retrain Equivalence in Machine Unlearning
- Does GenAI Rewrite How We Write? An Empirical Study on Two-Million Preprints
- The Burden of Interactive Alignment with Inconsistent Preferences
- Temporal Blindness in Multi-Turn LLM Agents: Misaligned Tool Use vs. Human Time Perception
- Debiasing Reward Models by Representation Learning with Guarantees
- Think Twice: Branch-and-Rethink Reasoning Reward Model
- Lightweight Robust Direct Preference Optimization
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- Rethinking Error: “Hallucinations” and Epistemological Indifference
- Neural language models as content analysis tools in psychology
- Larger and more instructable language models become less reliable
- POPI: Personalizing LLMs via Optimized Natural Language Preference Inference
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- Education Paradigm Shift To Maintain Human Competitive Advantage Over AI
- The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
- Large language model-based task planning for service robots: A review
- Adaptive Blockwise Search: Inference-Time Alignment for Large Language Models
- Code Aesthetics with Agentic Reward Feedback
- Smaller Models, Smarter Rewards: A Two-Sided Approach to Process and Outcome Rewards
- Can Language Models Compose Skills In-Context?
- MGFRec: Towards Reinforced Reasoning Recommendation with Multiple Groundings and Feedback
- Offline Preference Optimization via Maximum Marginal Likelihood Estimation
- Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust
- Retracing the Past: LLMs Emit Training Data When They Get Lost
- Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models
- FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
- Sentra-Guard: A Multilingual Human-AI Framework for Real-Time Defense Against Adversarial LLM Jailbreaks
- Aligning Diffusion Language Models via Unpaired Preference Optimization
- Towards Scalable Oversight via Partitioned Human Supervision
- Frustratingly Easy Task-aware Pruning for Large Language Models
- Agent-GSPO: Communication-Efficient Multi-Agent Systems via Group Sequence Policy Optimization
- Feature-Guided SAE Steering for Refusal-Rate Control using Contrasting Prompts
- Chitchat with AI: Understand the supply chain carbon disclosure of companies worldwide through Large Language Model
- A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
- GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
- PACR: Progressively Ascending Confidence Reward for LLM Reasoning
- You Don't Need Prompt Engineering Anymore: The Prompting Inversion
- DETECT: Determining Ease and Textual Clarity of German Text Simplifications
- OlaMind: Towards Human-Like and Hallucination-Safe Customer Service for Retrieval-Augmented Dialogue
- Controllable Mathematical Reasoning via Self-Optimizing Thought Vectors
- Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
- Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
- Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
- α-LoRA: Effective Fine-Tuning via Base Model Rescaling
- Weak-to-Strong Generalization under Distribution Shifts
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
- Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference
- Social Simulations with Large Language Model Risk Utopian Illusion
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generation
- Beyond Pairwise: Empowering LLM Alignment With Ranked Choice Modeling
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- Opening up ChatGPT: Tracking openness, transparency, and accountability in instruction-tuned text generators
- Exploring Parameter-Efficient Fine-Tuning Techniques for Code Generation with Large Language Models
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- Epipolar Geometry Improves Video Generation Models
- Learning Grouped Lattice Vector Quantizers for Low-Bit LLM Compression
- Compress to Impress: Efficient LLM Adaptation Using a Single Gradient Step on 100 Samples
- Shoot First, Ask Questions Later? Building Rational Agents that Explore and Act Like People
- Language Ranker: A Lightweight Ranking framework for LLM Decoding
- Robust Preference Alignment via Directional Neighborhood Consensus
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- BoundRL: Efficient Structured Text Segmentation through Reinforced Boundary Generation
- Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning
- Vox-Evaluator: Enhancing Stability and Fidelity for Zero-shot TTS with A Multi-Level Evaluator
- Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
- No-Regret Thompson Sampling for Finite-Horizon Markov Decision Processes with Gaussian Processes
- An Empirical Study of Sample Selection Strategies for Large Language Model Repair
- KL-Regularized Reinforcement Learning is Designed to Mode Collapse
- Black Box Absorption: LLMs Undermining Innovative Ideas
- RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging
- ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows
- Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
- g-DPO: Scalable Preference Optimization for Protein Language Models
- Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
- Enhancing visual-LLM for construction site safety compliance via prompt engineering and Bi-stage retrieval-augmented generation
- Dialogue Is Not Enough to Make a Communicative BabyLM (But Neither Is Developmentally Inspired Reinforcement Learning)
- Data-Centric Lessons To Improve Speech-Language Pretraining
- EQPO: Equitable Group Relative Policy Optimization for Clinical Reasoning
- Review of Tools for Zero-Code LLM Based Application Development
- SynCast: Synergizing Contradictions in Precipitation Nowcasting via Diffusion Sequential Preference Optimization
- PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models
- HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimization
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- No Compute Left Behind: Rethinking Reasoning and Sampling with Masked Diffusion Models
- RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
- The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- QKCV Attention: Enhancing Time Series Forecasting with Static Categorical Embeddings for Both Lightweight and Pre-trained Foundation Models
- Search Self-play: Pushing the Frontier of Agent Capability without Supervision
- KAT-Coder Technical Report
- Verifiable Accuracy and Abstention Rewards in Curriculum RL to Alleviate Lost-in-Conversation
- HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models
- Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
- Large language models in medicine
- Extracting alignment data in open models
- Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
- Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
- StarBench: A Turn-Based RPG Benchmark for Agentic Multimodal Decision-Making and Information Seeking
- Chain-of-Conceptual-Thought Elicits Daily Conversation in Large Language Models
- AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library
- ECG-LLM -- training and evaluation of domain-specific large language models for electrocardiography
- ADPO: Anchored Direct Preference Optimization
- ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control
- IF-VidCap: Can Video Caption Models Follow Instructions?
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- Counterfactual Reasoning for Steerable Pluralistic Value Alignment of Large Language Models
- Heterogeneous Adversarial Play in Interactive Environments
- DP2O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution
- Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning
- Mapping Post-Training Forgetting in Language Models at Scale
- Planned Diffusion
- Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models
- Assessing Monotone Dependence: Area Under the Curve Meets Rank Correlation
- Unbiased Gradient Low-Rank Projection
- Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
- A Principle of Targeted Intervention for Multi-Agent Reinforcement Learning
- Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
- A Mimamsa Inspired Framework For Instruction Sequencing In AI Agents
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
- Zero‐ and few‐shot prompting of generative large language models provides weak assessment of risk of bias in clinical trials
- Agentic Reinforcement Learning for Search Misaligns Instruction-Tuning
- Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
- The Atomic Instruction Gap: Instruction-Tuned LLMs Struggle with Simple, Self-Contained Directives
- Strengthening LLMs for Tabular Prediction with Structural Priors
- Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
- Intent-Driven LLM Ensemble Planning for Flexible Multi-Robot Disassembly: Demonstration on EV Batteries
- Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models
- Fine-tuning Flow Matching Generative Models with Intermediate Feedback
- Rewarding the Journey, Not Just the Destination: A Composite Path and Answer Self-Scoring Reward Mechanism for Test-Time Reinforcement Learning
- JT-Safe: Intrinsically Enhancing the Safety and Trustworthiness of LLMs
- Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
- OG-Rank: Learning to Rank Fast and Slow with Uncertainty and Reward-Trend Guided Adaptive Exploration
- Annotation-Efficient Universal Honesty Alignment
- Mapping from Meaning: Addressing the Miscalibration of Prompt-Sensitive Language Models
- SafeSearch: Do Not Trade Safety for Utility in LLM Search Agents
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- A Comprehensive Survey on Reinforcement Learning-based Agentic Search: Foundations, Roles, Optimizations, Evaluations, and Applications
- AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
- QuanBench: Benchmarking Quantum Code Generation with Large Language Models
- Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
- CRepair Wrapper Improves Structural Self-Repair Across Three LLM Families: A Cross-Model Replication Study
- Dual-Weighted Reinforcement Learning for Generative Preference Modeling
- ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
- InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training
- Voting with the Graph: Stable RLAIF via Topological Consistency Maximization
- The Road Less Traveled: Enhancing Exploration in LLMs via Sequential Sampling
- MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
- Direct Preference Optimization with Unobserved Preference Heterogeneity: The Necessity of Ternary Preferences
- Stochastic Optimization with Random Search
- MARSHAL: Incentivizing Multi-Agent Reasoning via Self-Play with Strategic LLMs
- STABLE: Gated Continual Learning for Large Language Models
- Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning
- Continual Learning via Sparse Memory Finetuning
- DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
- Reasoning with Sampling: Your Base Model is Smarter Than You Think
- Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
- Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
- Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
- Holdout-Loss-Based Data Selection for LLM Finetuning via In-Context Learning
- Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
- Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback
- Inference-Time Search Using Side Information for Diffusion-Based Image Reconstruction
- Your Next Token Prediction: A Multilingual Benchmark for Personalized Response Generation
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
- Stop-RAG: Value-Based Retrieval Control for Iterative RAG
- Large Reasoning Embedding Models: Towards Next-Generation Dense Retrieval Paradigm
- Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space
- Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography Segmentation
- Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
- Where to Search: Measure the Prior-Structured Search Space of LLM Agents
- Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning
- Budget-aware Test-time Scaling via Discriminative Verification
- Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
- Training LLM Agents to Empower Humans
- Confidence as a Reward: Transforming LLMs into Reward Models
- M2PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation
- Visual Interestingness Decoded: How GPT-4o Mirrors Human Interests
- Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanism
- Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling
- PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
- From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails
- Beyond Imitation: Recovering Dense Rewards from Demonstrations
- Improved Robustness of Deep Reinforcement Learning for Control of Time-Varying Systems by Bounded Extremum Seeking
- A11YN: aligning LLMs for accessible web UI code generation
- The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
- ConsintBench: Evaluating Language Models on Real-World Consumer Intent Understanding
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
- Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
- SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
- Automated Network Protocol Testing with LLM Agents
- Deliberate Lab: A Platform for Real-Time Human-AI Social Experiments
- Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
- Repairing Reward Functions with Human Feedback to Mitigate Reward Hacking
- On the Role of Preference Variance in Preference Optimization
- A Survey on Evaluation of Large Language Models
- How Well Can Preference Optimization Generalize Under Noisy Feedback?
- Data-Model Co-Evolution: Growing Test Sets to Refine LLM Behavior
- From Literal to Liberal: A Meta-Prompting Framework for Eliciting Human-Aligned Exception Handling in Large Language Models
- Expert or not? assessing data quality in offline reinforcement learning
- COSTAR-A: A prompting framework for enhancing Large Language Model performance on Point-of-View questions
- Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
- Laminar: A Scalable Asynchronous RL Post-Training Framework
- VISaGE: Understanding Visual Generics and Exceptions
- Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
- Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
- PromptFlow: Training Prompts Like Neural Networks
- Reinforced Preference Optimization for Recommendation
- ResearStudio: A Human-Intervenable Framework for Building Controllable Deep-Research Agents
- Self-Verifying Reflection Helps Transformers with CoT Reasoning
- Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editing
- Locket: Robust Feature-Locking Technique for Language Models
- Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback
- Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
- Hierarchical Alignment: Surgical Fine-Tuning via Functional Layer Specialization in Large Language Models
- CPR: Mitigating Large Language Model Hallucinations with Curative Prompt Refinement
- Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
- ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution
- Pruning Cannot Hurt Robustness: Certified Trade-offs in Reinforcement Learning
- EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus
- LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
- Don't Walk the Line: Boundary Guidance for Filtered Generation
- Analyzing and Internalizing Complex Policy Documents for LLM Agents
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
- Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
- Valid Survey Simulations with Limited Human Data: The Roles of Prompting, Fine-Tuning, and Rectification
- DocReward: A Document Reward Model for Structuring and Stylizing
- Vision-LLMs for Spatiotemporal Traffic Forecasting
- Exploring and Leveraging Class Vectors for Classifier Editing
- AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
- Automating Structural Engineering Workflows with Large Language Model Agents
- APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
- Rediscovering Entropy Regularization: Adaptive Coefficient Unlocks Its Potential for LLM Reinforcement Learning
- Cog-Rethinker: Hierarchical Metacognitive Reinforcement Learning for LLM Reasoning
- BanglaMATH : A Bangla benchmark dataset for testing LLM mathematical reasoning at grades 6, 7, and 8
- Enhancing Large Language Model Reasoning via Selective Critical Token Fine-Tuning
- Does LLM Focus on the Right Words? Mitigating Context Bias in LLM-based Recommenders
- AMiD: Knowledge Distillation for LLMs with α-mixture Assistant Distribution
- Direct Multi-Token Decoding
- Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
- Learning Dynamics of VLM Finetuning
- T-T: Table Transformer for Tagging-based Aspect Sentiment Triplet Extraction
- UpSafe^∘C: Upcycling for Controllable Safety in Large Language Models
- Understanding Sampler Stochasticity in Training Diffusion Models for RLHF
- Controllable Generative Trajectory Prediction via Weak Preference Alignment
- Hierarchical Optimization via LLM-Guided Objective Evolution for Mobility-on-Demand Systems
- DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
- Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
- MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
- PrediQL: Automated Testing of GraphQL APIs with LLMs
- VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
- OpusAnimation: Code-Based Dynamic Chart Generation
- ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test
- MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
- Reasoning-Enhanced Large Language Models for Molecular Property Prediction
- You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
- PIXEL: Adaptive Steering Via Position-wise Injection with eXact Estimated Levels under Subspace Calibration
- CompassNav: Steering From Path Imitation To Decision Understanding In Navigation
- Breaking the Likelihood Trap: Consistent Generative Recommendation with Graph-structured Model
- Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models
- Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning
- Beyond the limitation of a single query: Train your LLM for query expansion with Reinforcement Learning
- RIPRAG: Hack a Black-box Retrieval-Augmented Generation Question-Answering System with Reinforcement Learning
- Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
- A-IPO: Adaptive Intent-driven Preference Optimization
- Artificial intelligence as a surrogate brain: Bridging neural dynamical models and data
- Contemplative Superalignment
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
- Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
- Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplaces
- On the Role of Domain Experts in Creating Effective Tutoring Systems
- Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
- IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
- Signals: Trajectory Sampling and Triage for Agentic Interactions
- Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
- An Alternative Trajectory for Generative AI
- Can RL Improve Generalization of LLM Agents? An Empirical Study
- Don't Throw Away Your Pretrained Model
- Token Is All You Price
- SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models
- ConDABench: Interactive Evaluation of Language Models for Data Analysis
- Agentic Systems in Radiology: Design, Applications, Evaluation, and Challenges
- Safety Game: Balancing Safe and Informative Conversations with Blackbox Agentic AI using LP Solvers
- DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
- GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis
- DSPO: Stable and Efficient Policy Optimization for Agentic Search and Reasoning
- AdaPM: a Partial Momentum Algorithm for LLM Training
- Student Development Agent: Risk-free Simulation for Evaluating AIED Innovations
- Users as Annotators: LLM Preference Learning from Comparison Mode
- Leading the Follower: Learning Persuasive Agents in Social Deduction Games
- Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
- SHERLOCK: Towards Dynamic Knowledge Adaptation in LLM-enhanced E-commerce Risk Management
- SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
- Score-Based Density Estimation from Pairwise Comparisons
- Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Markov Likelihood
- Active Model Selection for Large Language Models
- KORMo: Korean Open Reasoning Model for Everyone
- Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models
- VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search
- Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
- From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP
- Opponent Shaping in LLM Agents
- LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
- Memory Retrieval and Consolidation in Large Language Models through Function Tokens
- Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
- Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
- AILoRA: Function-Aware Asymmetric Initialization for Low-Rank Adaptation of Large Language Models
- LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?
- Contrastive Weak-to-strong Generalization
- Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
- MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
- Energy-Driven Steering: Reducing False Refusals in Large Language Models
- Stop DDoS Attacking the Research Community with AI-Generated Survey Papers
- Dynamic Generation of Multi-LLM Agents Communication Topologies with Graph Diffusion Models
- GCPO: When Contrast Fails, Go Gold
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
- Next-Generation LLM for UAV: From Natural Language to Autonomous Flight
- Mix- and MoE-DPO: A Variational Inference Approach to Direct Preference Optimization
- Efficient Preference-Based Reinforcement Learning: Randomized Exploration Meets Experimental Design
- Position: Privacy Is Not Just Memorization!
- AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
- VisualDAN: Exposing Vulnerabilities in VLMs with Visual-Driven DAN Commands
- On the optimization dynamics of RLVR: Gradient gap and step size thresholds
- Drift No More? Context Equilibria in Multi-Turn LLM Interactions
- MLLM as a UI Judge: Benchmarking Multimodal LLMs for Predicting Human Perception of User Interfaces
- Self-Improving LLM Agents at Test-Time
- Reinforcing Diffusion Models by Direct Group Preference Optimization
- LiveThinking: Enabling Real-Time Efficient Reasoning for AI-Powered Livestreaming via Reinforcement Learning
- Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy Optimization
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learning
- An Adaptive Multi Agent Bitcoin Trading System
- FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
- TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
- Prepared mind, fast response: A temporal decoupling framework for adaptive knowledge orchestration in open-domain dialogue
- Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
- Post-Norm can Resharpen Attention
- From Data to Rewards: a Bilevel Optimization Perspective on Maximum Likelihood Estimation
- Phase Diagram of Dropout for Two-Layer Neural Networks in the Mean-Field Regime
- MAPRO: Recasting Multi-Agent Prompt Optimization as Maximum a Posteriori Inference
- On the Convergence of Moral Self-Correction in Large Language Models
- Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- Reasoning for Hierarchical Text Classification: The Case of Patents
- AI for Abolition? A Participatory Design Approach
- LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish
- Prompt Optimization Across Multiple Agents for Representing Diverse Human Populations
- Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
- LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling
- Prototyping Multimodal GenAI Real-Time Agents with Counterfactual Replays and Hybrid Wizard-of-Oz
- AMAS: Adaptively Determining Communication Topology for LLM-based Multi-Agent System
- Experiential Reinforcement Learning
- Authenticated Workflows: A Systems Approach to Protecting Agentic AI
- CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
- Generating Meaning: Active Inference and the Scope and Limits of Passive AI
- Predictive Preference Learning from Human Interventions
- DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning
- Do LLMs Know They Are Being Tested? Evaluation Awareness and Incentive-Sensitive Failures in GPT-OSS-20B
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- GRACE: A Language Model Framework for Explainable Inverse Reinforcement Learning
- Agent-in-the-Loop: A Data Flywheel for Continuous Improvement in LLM-based Customer Support
- ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
- Aligning Large Language Models via Fully Self-Synthetic Data
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization
- Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation
- Optimal Stopping vs Best-of-N for Inference Time Optimization
- Online Rubrics Elicitation from Pairwise Comparisons
- Reasoning by Exploration: A Unified Approach to Retrieval and Generation over Graphs
- POME: Post Optimization Model Edit via Muon-style Projection
- Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
- EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
- LLM Bias Detection and Mitigation through the Lens of Desired Distributions
- Taxonomy of User Needs and Actions
- The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
- Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL
- EARL: Efficient Agentic Reinforcement Learning Systems for Large Language Models
- Prompt reinforcing for long-term planning of large language models
- Optimizing for Persuasion Improves LLM Generalization: Evidence from Quality-Diversity Evolution of Debate Strategies
- EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
- Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
- Provably Convergent Primal-Dual DPO for Constrained LLM Alignment
- On the Role of Difficult Prompts in Self-Play Preference Optimization
- Vul-R2: A Reasoning LLM for Automated Vulnerability Repair
- Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
- When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
- Classical AI vs. LLMs for Decision-Maker Alignment in Health Insurance Choices
- MADIAVE: Multi-Agent Debate for Implicit Attribute Value Extraction
- Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique
- Prototype-Based Dynamic Steering for Large Language Models
- Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
- The Answer Lies Within: Self-Derived Rewards Enable Explainable Relation Extraction
- Bloom: Designing for LLM-Augmented Behavior Change Interactions
- AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering
- InvThink: Premortem Reasoning for Safer Language Models
- Beyond Monolithic Rewards: A Hybrid and Multi-Aspect Reward Optimization for MLLM Alignment
- Chrysalis: A Unified System for Comparing Active Teaching and Passive Learning with AI Agents in Education
- Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
- RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
- From Noisy Traces to Stable Gradients: Bias-Variance Optimized Preference Optimization for Aligning Large Reasoning Models
- TeachLM: Post-Training LLMs for Education Using Authentic Learning Data
- Alignment Tipping Process: How Self-Evolution Pushes LLM Agents Off the Rails
- Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning
- Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
- FT-MDT: Extracting Decision Trees from Medical Texts via a Novel Low-rank Adaptation Method
- EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents
- FedSRD: Sparsify-Reconstruct-Decompose for Communication-Efficient Federated Large Language Models Fine-Tuning
- Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
- Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
- VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- MedCLM: Learning to Localize and Reason via a CoT-Curriculum in Medical Vision-Language Models
- DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization
- Good Intentions Beyond ACL: Who Does NLP for Social Good, and Where?
- Proactive defense against LLM Jailbreak
- COLE: a Comprehensive Benchmark for French Language Understanding Evaluation
- Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
- GRACE: Generative Representation Learning via Contrastive Policy Optimization
- VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
- LLM Based Bayesian Optimization for Prompt Search
- Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
- Self-supervised diffusion model fine-tuning for costate initialization using Markov chain Monte Carlo
- Spatiotemporal Forecasting as Planning: A Model-Based Reinforcement Learning Approach with Generative World Models
- Reflection Before Action: Designing a Framework for Quantifying Thought Patterns for Increased Self-awareness in Personal Decision Making
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment frm Heterogeneous Rewards
- Thinking on the Fly: Test-Time Reasoning Enhancement via Latent Thought Policy Optimization
- Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
- Fine-Grained GRPO for Precise Preference Alignment in Flow Models
- Large Language Models Hallucination: A Comprehensive Survey
- Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
- Increasing LLM response trustworthiness using voting ensembles
- Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- The Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic View
- Activation Steering with a Feedback Controller
- JEF-Hinter: Leveraging Offline Knowledge for Improving Web Agents Adaptation
- Equipping Retrieval-Augmented Large Language Models with Document Structure Awareness
- AgentRL: Scaling Agentic Reinforcement Learning with a Multi-Turn, Multi-Task Framework
- MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
- RESTRAIN: From Spurious Votes to Signals -- Self-Driven RL with Self-Penalization
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation
- How Catastrophic is Your LLM? Certifying Risk in Conversation
- Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
- TROLL: Trust Regions improve Reinforcement Learning for Large Language Models
- OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
- TreePrompt: Leveraging Hierarchical Few-Shot Example Selection for Improved English-Persian and English-German Translation
- Group Policy Gradient
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learning
- Decoupling Task-Solving and Output Formatting in LLM Generation
- Generalized Fitted Q-Iteration with Clustered Data
- Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
- Triplet-Structured Knowledge Integration for Multi-Turn Medical Reasoning
- Best-of-Majority: Minimax-Optimal Strategy for Pass@k Inference Scaling
- RoiRL: Efficient, Self-Supervised Reasoning with Offline Iterative Reinforcement Learning
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs
- A Granular Study of Safety Pretraining under Model Abliteration
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
- Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- Self-Improvement in Multimodal Large Language Models: A Survey
- Smart-GRPO: Smartly Sampling Noise for Efficient RL of Flow-Matching Models
- Fine-Tuning on Noisy Instructions: Effects on Generalization and Performance
- Fine-Tuning Jailbreaks under Highly Constrained Black-Box Settings: A Three-Pronged Approach
- Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
- GRAD: Generative Retrieval-Aligned Demonstration Sampler for Efficient Few-Shot Reasoning
- Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?
- mR3: Multilingual Rubric-Agnostic Reward Reasoning Models
- Strategic Fusion of Vision Language Models: Shapley-Credited Context-Aware Dawid-Skene for Multi-Label Tasks in Autonomous Driving
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Multi-Actor Multi-Critic Deep Deterministic Reinforcement Learning with a Novel Q-Ensemble Method
- GEM: A Gym for Agentic LLMs
- Uncovering the Computational Ingredients of Human-Like Representations in LLMs
- It Takes Two: Your GRPO Is Secretly DPO
- From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
- Inclusive Easy-to-Read Generation for Individuals with Cognitive Impairments
- ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewards
- On Predictability of Reinforcement Learning Dynamics for Large Language Models
- Rethinking Reward Models for Multi-Domain Test-Time Scaling
- Structuring Reasoning for Complex Rules Beyond Flat Representations
- A Call to Action for a Secure-by-Design Generative AI Paradigm
- Train on Validation (ToV): Fast data selection with applications to fine-tuning
- Retrieval-Augmented Framework for LLM-Based Clinical Decision Support
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- Making, not Taking, the Best of N
- BroRL: Scaling Reinforcement Learning via Broadened Exploration
- Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning across the Model Capability Continuum
- Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
- Stabilizing Policy Gradients for Sample-Efficient Reinforcement Learning in LLM Reasoning
- Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
- Prompt Curriculum Learning for Efficient LLM Post-Training
- Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- Judging with Confidence: Calibrating Autoraters to Preference Distributions
- Debunk the Myth of SFT Generalization
- GRPO-λ: Credit Assignment improves LLM Reasoning
- Thinkquel: A Model Dedicated to Text-to-dbt Using Synthetic Data and a Span-Aware Objective
- PrimeX: A Dataset of Worldview, Opinion, and Explanation
- Which Rewards Matter? Reward Selection for Reinforcement Learning under Limited Feedback
- Recursive Self-Aggregation Unlocks Deep Thinking in Large Language Models
- Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning
- SafeBehavior: Simulating Human-Like Multistage Reasoning to Mitigate Jailbreak Attacks in Large Language Models
- One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient
- Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
- Alignment-Aware Decoding
- DyFlow: Dynamic Workflow Framework for Agentic Reasoning
- RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation
- RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- Efficient On-Policy Reinforcement Learning via Exploration of Sparse Parameter Space
- Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Supporting Creative Ownership through Deep Learning-Based Music Variation
- Improving Sampling Efficiency in RLVR through Adaptive Rollout and Response Reuse
- Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking
- Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
- OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
- MuPlon: Multi-Path Causal Optimization for Claim Verification through Controlling Confounding
- The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- PCPO: Proportionate Credit Policy Optimization for Aligning Image Generation Models
- Learning to Reason as Action Abstractions with Scalable Mid-Training RL
- Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
- RFG: Test-Time Scaling for Diffusion Large Language Model Reasoning with Reward-Free Guidance
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- IRIS: Intrinsic Reward Image Synthesis
- Aligning Multilingual Reasoning with Verifiable Semantics from a High-Resource Expert Model
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
- Toxicity in Online Platforms and AI Systems: A Survey of Needs, Challenges, Mitigations, and Future Directions
- A Method for Quantifying Human Risk and a Blueprint for LLM Integration
- Fingerprinting LLMs via Prompt Injection
- Polychromic Objectives for Reinforcement Learning
- Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
- UniAPL: A Unified Adversarial Preference Learning Framework for Instruct-Following
- The Era of Real-World Human Interaction: RL from User Conversations
- Rethinking Entropy Regularization in Large Reasoning Models
- CLPO: Curriculum Learning meets Policy Optimization for LLM Reasoning
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- SecInfer: Preventing Prompt Injection via Inference-time Scaling
- Learning Distinguishable Representations in Deep Q-Networks for Linear Transfer
- Retro*: Optimizing LLMs for Reasoning-Intensive Document Retrieval
- T-POP: Test-Time Personalization with Online Preference Feedback
- Reference-Free Rating of LLM Responses via Latent Information
- CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task
- Multilingual Text-to-SQL: Benchmarking the Limits of Language Models with Collaborative Language Agents
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
- PEARL: Performance-Enhanced Aggregated Representation Learning
- DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
- Graph Optimization Foundation Model: Tokenizing Graph via A Language-Model Paradigm
- Prompt and Parameter Co-Optimization for Large Language Models
- Humanline: Online Alignment as Perceptual Loss
- Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
- RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
- World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
- The problem of alignment
- Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
- Interactive Groupwise Comparison for Reinforcement Learning from Human Feedback
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play
- Agentar-Scale-SQL: Advancing Text-to-SQL through Orchestrated Test-Time Scaling
- Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs
- Reinforcement Mid-Training
- Preference-Based Dynamic Ranking Structure Recognition
- Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE
- EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
- ADAPT: Lightweight, Long-Range Machine Learning Force Fields Without Graphs
- Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training
- MACE: A Hybrid LLM Serving System with Colocated SLO-aware Continuous Retraining Alignment
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- Reinforcement Learning with Inverse Rewards for World Model Post-training
- Winning the Pruning Gamble: A Unified Approach to Joint Sample and Token Pruning for Efficient Supervised Fine-Tuning
- Rethinking Reward Miscalibration of GRPO in Agentic RL
- Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules
- Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions
- Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
- Anchored Supervised Fine-Tuning
- Why Alignment Must Precede Distillation: A Minimal Working Explanation
- Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs
- Fast Thinking for Large Language Models
- Towards Efficient CoT Distillation: Self-Guided Rationale Selector for Better Performance with Fewer Rationales
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
- Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
- On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
- Large-Scale Constraint Generation -- Can LLMs Parse Hundreds of Constraints?
- Advancing Multi-agent Traffic Simulation via R1-Style Reinforcement Fine-Tuning
- EAPO: Enhancing Policy Optimization with On-Demand Expert Assistance
- Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models
- Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
- Cognition-of-Thought Elicits Social-Aligned Reasoning in Large Language Models
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
- MedCritical: Enhancing Medical Reasoning in Small Language Models via Self-Collaborative Correction
- Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
- Decoupling Reasoning and Perception: An LLM-LMM Framework for Faithful Visual Reasoning
- Scaling Policy Compliance Assessment in Language Models with Policy Reasoning Traces
- A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Knowledge distillation through geometry-aware representational alignment
- Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts
- Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs
- General Exploratory Bonus for Optimistic Exploration in RLHF
- Multiplayer Nash Preference Optimization
- DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning
- Risk Profiling and Modulation for LLMs
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
- WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning
- Adaptive Margin RLHF via Preference over Preferences
- Causally-Enhanced Reinforcement Policy Optimization
- Reinforcement Learning-Based Prompt Template Stealing for Text-to-Image Models
- MTRec: Learning to Align with User Preferences via Mental Reward Models
- ChatGPT for complex text evaluation tasks
- Voice user interfaces for effortless navigation in medical virtual reality environments
- A shared model-based linguistic space for transmitting our thoughts from brain to brain in natural conversations
- Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
- Optimizing Long-Form Clinical Text Generation with Claim-Based Rewards
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
- Boosting Pointer Analysis With LLM-Enhanced Allocation Function Detection
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- RAPID3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
- Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models
- When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance
- Context Parametrization with Compositional Adapters
- Multilingual Vision-Language Models, A Survey
- S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
- The Rogue Scalpel: Activation Steering Compromises LLM Safety
- Goal-Guided Efficient Exploration via Large Language Model in Reinforcement Learning
- Black-Box Hallucination Detection via Consistency Under the Uncertain Expression
- Exposing Hallucinations To Suppress Them: VLMs Representation Editing With Generative Anchors
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards
- Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
- Synthetic Dialogue Generation for Interactive Conversational Elicitation & Recommendation (ICER)
- SoK: Potentials and Challenges of Large Language Models for Reverse Engineering
- Can LLMs Solve and Generate Linguistic Olympiad Puzzles?
- FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
- ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
- Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
- AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
- Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
- In-Context Learning can Perform Continual Learning Like Humans
- Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective
- Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization
- We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
- Compute-Optimal Quantization-Aware Training
- RLP: Reinforcement as a Pretraining Objective
- Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models
- MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning
- DriftLite: Lightweight Drift Control for Inference-Time Scaling of Diffusion Models
- Multi-Objective Reinforcement Learning for Large Language Model Optimization: Visionary Perspective
- Hallucination reduction with CASAL: Contrastive Activation Steering For Amortized Learning
- Correct Reasoning Paths Visit Shared Decision Pivots
- Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain Generation
- Preemptive Detection and Steering of LLM Misalignment via Latent Reachability
- Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
- Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
- The role of synthetic data in Multilingual, Multi-cultural AI systems: Lessons from Indic Languages
- SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
- LLMTrace: A Corpus for Classification and Fine-Grained Localization of AI-Written Text
- Fine-tuning of Large Language Models for Domain-Specific Cybersecurity Knowledge
- LogReasoner: Empowering LLMs with Expert-like Coarse-to-Fine Reasoning for Automated Log Analysis
- Physics of Learning: A Lagrangian perspective to different learning paradigms
- CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning
- It's Not You, It's Clipping: A Soft Trust-Region via Probability Smoothing for LLM RL
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning
- Difference-Guided Reasoning: A Temporal-Spatial Framework for Large Language Models
- Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
- PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
- Actor-Critic without Actor
- Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
- d2: Improving Reasoning in Diffusion Language Models via Trajectory Likelihood Estimation
- Enhancing Python Programming Education with an AI-Powered Code Helper: Design, Implementation, and Impact
- Complexity-Regularized Proximal Policy Optimization
- Instruction Boundary: Quantifying Biases in LLM Reasoning under Various Coverage
- ToolBrain: A Flexible Reinforcement Learning Framework for Agentic Tools
- Failure Modes of Maximum Entropy RLHF
- OR-Toolformer: Modeling and Solving Operations Research Problems with Tool Augmented Large Language Models
- Federation of Agents: A Semantics-Aware Communication Fabric for Large-Scale Agentic AI
- V-GameGym: Visual Game Generation for Code Large Language Models
- LatentGuard: Controllable Latent Steering for Robust Refusal of Attacks and Reliable Response Generation
- Embodied AI: From LLMs to World Models
- The Knowledge-Behaviour Disconnect in LLM-based Chatbots
- MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
- WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
- PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
- Future Policy Aware Preference Learning for Mathematical Reasoning
- PolicyPad: Collaborative Prototyping of LLM Policies
- Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation
- Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
- UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
- bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
- ExPe: Exact Positional Encodings for Generative Transformer Models with Extrapolating Capabilities
- From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
- Training Skills Like Parameters via Self-Supervised Semantic Diffusion
- What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
- RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
- DeepResearch Agent System
- Compliance2LoRA: Personalizable On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters
- ACPO: Asymmetric Credit Policy Optimization via Mode-Local Entropy Surrogate
- MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
- FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
- LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models
- TAPO: Transition-Aware Policy Optimization for LLM Agents
- MMLDSum-LLM: Multimodal Long-Document Summarization with Visual-Alignment and Keyword-Aware
- Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
- IFHierBench: Hierarchical Instruction Following for Large Language Models
- Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
- Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models
- SDO: Structure-Aware Data Organization for Efficient LLM Post-Training
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
- RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
- MedLLM: An Open Medical Language Model at the Sub-Billion Scale
- ToolRec: Calibrated Preference Alignment for Query Recommendation in On-Device Assistants
- Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
- Divergence Decoding: Training-Free Capability Fusion
- Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
- Machine learning in computational literary studies
- A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing
- DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes
- Implicit Safety Alignment from Crowd Preferences
- Persuading large language models to comply with objectionable requests
- Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
- S-GRPO: Unified Post-Training for Large Vision-Language Models
- Paper Espresso: From Paper Overload to Research Insight
- RELISH: LLM REgression with a Latent Iterative State Head
- GenAI and the Mirage of Personalised Learning for All
- From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency
- Game-theory behaviour of large language models: The case of Keynesian beauty contests
- A Theory of Appropriateness That Accounts for Norms of Rationality
- A large language model-based agent for wayfinding: simulation of spatial perception and memory
- Pressure Reveals Character: Behavioural Alignment Evaluation at Depth
- References Improve LLM Alignment in Non-Verifiable Domains
- SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
- Assessing and alleviating state anxiety in large language models
- Security and Privacy Challenges of Large Language Models: A Survey
- Auditing large language models: a three-layered approach
- Assessing political bias in AI systems: a framework for disentangling viewpoint preferences from epistemic integrity
- Bypassing Guardrails: Lessons Learned from Red Teaming ChatGPT
- Toward Human-Centered Explainability: Natural Language Explanations for Anomaly Detection
- GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
- Summary of ChatGPT-Related research and perspective towards the future of large language models
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
- Reinforcement Learning on Pre-Training Data
- Agentic Reinforcement Learning with Implicit Step Rewards
- Speculative Safety-Aware Decoding
- SMITE: Enhancing Fairness in LLMs through Optimal In-Context Example Selection via Dynamic Validation
- Central Limit Theorems for Asynchronous Averaged Q-Learning
- Direct Preference Optimization for Speech Autoregressive Diffusion Models
- Diversity Boosts AI-Generated Text Detection
- MAPO: Mixed Advantage Policy Optimization
- SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer
- Enhancing LLM-Based Social Bot via an Adversarial Learning Framework
- PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generation
- LLMZ+: Contextual Prompt Whitelist Principles for Agentic LLMs
- APRIL: Active Partial Rollouts in Reinforcement Learning to Tame Long-tail Generation
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- NGRPO: Negative-enhanced Group Relative Policy Optimization
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- Advances in Large Language Models for Medicine
- GRPO++: Enhancing Dermatological Reasoning under Low Resource Settings
- Confidence-Aware Routing for Large Language Model Reliability Enhancement: A Multi-Signal Approach to Pre-Generation Hallucination Mitigation
- When Meaning Stays the Same, but Models Drift: Evaluating Quality of Service under Token-Level Behavioral Instability in LLMs
- The Narcissus Hypothesis: Descending to the Rung of Illusion
- ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- Automated Knowledge Graph Construction using Large Language Models and Sentence Complexity Modelling
- Exploiting Tree Structure for Credit Assignment in RL Training of LLMs
- Weights-Rotated Preference Optimization for Large Language Models
- LAD-VF: LLM-Automatic Differentiation Enables Fine-Tuning-Free Robot Planning from Formal Methods Feedback
- An Artificial Intelligence Value at Risk Approach: Metrics and Models
- ATLAS: Benchmarking and Adapting LLMs for Global Trade via Harmonized Tariff Code Classification
- A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialogue
- PG-CE: A Progressive Generation Dataset with Constraint Enhancement for Controllable Text Generation
- Understanding Post-Training Structural Changes in Large Language Models
- DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
- nDNA -- the Semantic Helix of Artificial Cognition
- IDfRA: Self-Verification for Iterative Design in Robotic Assembly
- Advancing Speech Understanding in Speech-Aware Language Models with GRPO
- Preference Distillation via Value based Reinforcement Learning
- SemanticGarment: Semantic-Controlled Generation and Editing of 3D Gaussian Garments
- Can GRPO Boost Complex Multimodal Table Understanding?
- LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
- SFT-TA: Supervised Fine-Tuned Agents in Multi-Agent LLMs for Automated Inductive Thematic Analysis
- USB-Rec: An Effective Framework for Improving Conversational Recommendation Capability of Large Language Model
- Improving User Interface Generation Models from Designer Feedback
- Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories
- Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models
- From Uniform to Heterogeneous: Tailoring Policy Optimization to Every Token's Nature
- RLGF: Reinforcement Learning with Geometric Feedback for Autonomous Driving Video Generation
- Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
- Assessing Classical Machine Learning and Transformer-based Approaches for Detecting AI-Generated Research Text
- The Programmer’s Assistant: Conversational Interaction with a Large Language Model for Software Development
- RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
- Near-Optimal Sample Complexity Bounds for Constrained Average-Reward MDPs
- SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning
- The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
- Redefining Experts: Interpretable Decomposition of Language Models for Toxicity Mitigation
- Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection
- The Alignment Bottleneck
- Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- REFER: Mitigating Bias in Opinion Summarisation via Frequency Framed Prompting
- Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
- The governance & behavioral challenges of generative artificial intelligence’s hypercustomization capabilities
- LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
- Reward Hacking Mitigation using Verifiable Composite Rewards
- How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
- Jamendo-QA: A Large-Scale Music Question Answering Dataset
- Dynamic Classifier-Free Diffusion Guidance via Online Feedback
- BaseReward: A Strong Baseline for Multimodal Reward Model
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models
- Generalizability of Large Language Model-Based Agents: A Comprehensive Survey
- RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
- Low-bit Model Quantization for Deep Neural Networks: A Survey
- Real, Fake, or Manipulated? Detecting Machine-Influenced Text
- Self-Improving Embodied Foundation Models
- AutoEdit: Automatic Hyperparameter Tuning for Image Editing
- CARGO: A Framework for Confidence-Aware Routing of Large Language Models
- Rationality Check! Benchmarking the Rationality of Large Language Models
- LLM Jailbreak Detection for (Almost) Free!
- Fleming-R1: Toward Expert-Level Medical Reasoning via Reinforcement Learning
- TDRM: Smooth Reward Models with Temporal Difference for LLM RL and Inference
- Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
- Assessing Historical Structural Oppression Worldwide via Rule-Guided Prompting of Large Language Models
- CodeFuse-CR-Bench: A Comprehensiveness-aware Benchmark for End-to-End Code Review Evaluation in Python Projects
- Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
- Synthetic bootstrapped pretraining
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- Dual-Actor Fine-Tuning of VLA Models: A Talk-and-Tweak Human-in-the-Loop Approach
- SAIL-VL2 Technical Report
- Learning the natural history of human disease with generative transformers
- An LLM-based multi-agent framework for agile effort estimation
- DSCC-HS: A Dynamic Self-Reinforcing Framework for Hallucination Suppression in Large Language Models
- Improving Context Fidelity via Native Retrieval-Augmented Reasoning
- DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
- Programmable Cognitive Bias in Social Agents
- A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks
- SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
- A deep reinforcement learning platform for antibiotic discovery
- Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews
- RepIt: Steering Language Models with Concept-Specific Refusal Vectors
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- The Anatomy of Alignment: Decomposing Preference Optimization by Steering Sparse Features
- Participatory AI: A Scandinavian Approach to Human-Centered AI
- Don't Change My View: Ideological Bias Auditing in Large Language Models
- Yet Another Watermark for Large Language Models
- Gender-Neutral Rewriting in Italian: Models, Approaches, and Trade-offs
- Shaping Explanations: Semantic Reward Modeling with Encoder-Only Transformers for GRPO
- Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
- Multi-Metric Preference Alignment for Generative Speech Restoration
- SSFO: Self-Supervised Faithfulness Optimization for Retrieval-Augmented Generation
- Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
- Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
- Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
- Active Domain Knowledge Acquisition with 100-Dollar Budget: Enhancing LLMs via Cost-Efficient, Expert-Involved Interaction in Sensitive Domains
- Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
- Scaling Agents via Continual Pre-training
- Human resource management in the age of generative artificial intelligence: Perspectives and research directions on ChatGPT
- Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
- Prompt Commons: Collective Prompting as Governance for Urban AI
- Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
- GTA: Supervised-Guided Reinforcement Learning for Text Classification with Large Language Models
- Exploring Conversational Design Choices in LLMs for Pedagogical Purposes: Socratic and Narrative Approaches for Improving Instructor's Teaching Practice
- MusicSwarm: Biologically Inspired Intelligence for Music Composition
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- POT: Inducing Overthinking in LLMs via Black-Box Iterative Optimization
- Pluralistic Off-policy Evaluation and Alignment
- Audited Reasoning Refinement: Fine-Tuning Language Models via LLM-Guided Step-Wise Evaluation and Correction
- When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- What Matters in Data for DPO?
- ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
- Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
- Pathological Truth Bias in Vision-Language Models
- ICR-RL: Deep Reinforcement Learning via In-Context Regression
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- Beyond IVR Touch-Tones: Customer Intent Routing using LLMs
- Continually Adding New Languages to Multilingual Language Models
- Teaching LLMs to Plan: Logical Chain-of-Thought Instruction Tuning for Symbolic Planning
- A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment
- Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning
- AI-Generated Content in Cross-Domain Applications: Research Trends, Challenges and Propositions
- When Are Two RLHF Objectives the Same?
- Limitations of refinement methods for weak to strong generalization
- The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language Models
- RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems
- SearchInstruct: Enhancing Domain Adaptation via Retrieval-Based Instruction Dataset Creation
- CrunchLLM: Multitask LLMs for Structured Business Reasoning and Outcome Prediction
- A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
- Pluralistic Alignment for Healthcare: A Role-Driven Framework
- Safety and Security Analysis of Large Language Models: Benchmarking Risk Profile and Harm Potential
- DeepDive: Advancing Deep Search Agents with Knowledge Graphs and Multi-Turn RL
- KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
- DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillation
- WebSight: A Vision-First Architecture for Robust Web Agents
- VARCO-VISION-2.0 Technical Report
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Multi-Intent Recognition in Dialogue Understanding: A Comparison Between Smaller Open-Source LLMs
- InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
- Towards Understanding Visual Grounding in Visual Language Models
- Inpainting-Guided Policy Optimization for Diffusion Large Language Models
- Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration
- Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented Tasks
- Topic-Guided Reinforcement Learning with LLMs for Enhancing Multi-Document Summarization
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining
- Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
- Latency and Token-Aware Test-Time Compute
- How well can LLMs provide planning feedback in grounded environments?
- Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling
- RewardDance: Reward Scaling in Visual Generation
- X-Teaming Evolutionary M2S: Automated Discovery of Multi-turn to Single-turn Jailbreak Templates
- Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing
- Generative Data Refinement: Just Ask for Better Data
- CM-Align: Consistency-based Multilingual Alignment for Large Language Models
- AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- Ubiquitous Intelligence Via Wireless Network-Driven LLMs Evolution
- Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- Language Self-Play For Data-Free Training
- TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
- The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward
- Uncovering Scaling Laws for Large Language Models via Inverse Problems
- EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention
- Fuzz4All: Universal Fuzzing with Large Language Models
- Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
- SoK: Security and Privacy of AI Agents for Blockchain
- Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
- UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward
- MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security
- Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
- Reinforcement Learning Foundations for Deep Research Systems: A Survey
- Probabilistic Modeling of Latent Agentic Substructures in Deep Neural Networks
- Outcome-based Exploration for LLM Reasoning
- Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation
- A Fragile Number Sense: Probing the Elemental Limits of Numerical Reasoning in LLMs
- The Thinking Therapist: Training Large Language Models to Deliver Acceptance and Commitment Therapy using Supervised Fine-Tuning and Odds Ratio Policy Optimization
- RL Fine-Tuning Heals OOD Forgetting in SFT
- Sovereign AI for 6G: Towards the Future of AI-Native Networks
- Uncovering the Vulnerability of Large Language Models in the Financial Domain via Risk Concealment
- EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
- Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint
- IntrEx: A Dataset for Modeling Engagement in Educational Conversations
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers
- Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting
- BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
- Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL
- Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
- Benchmarking Gender and Political Bias in Large Language Models
- Reverse-Engineered Reasoning for Open-Ended Generation
- Finetuning LLMs for Human Behavior Prediction in Social Science Experiments
- Understanding the Influence of Synthetic Data for Text Embedders
- Coefficients-Preserving Sampling for Reinforcement Learning with Flow Matching
- RL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMs
- Chatbot To Help Patients Understand Their Health
- CC-GSEO-Bench: A Content-Centric Benchmark for Measuring Source Influence in Generative Search Engines
- Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation
- ZhiFangDanTai: Fine-tuning Graph-based Retrieval-Augmented Generation Model for Traditional Chinese Medicine Formula
- Self-Aligned Reward: Towards Effective and Efficient Reasoners
- CURE: Controlled Unlearning for Robust Embeddings -- Mitigating Conceptual Shortcuts in Pre-Trained Language Models
- LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
- Cloning a Conversational Voice AI Agent from Call Recording Datasets for Telesales
- What-If Analysis of Large Language Models: Explore the Game World Using Proactive Thinking
- Post-training Large Language Models for Diverse High-Quality Responses
- A Lightweight Framework for Trigger-Guided LoRA-Based Self-Adaptation in LLMs
- Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
- Symbolic Graphics Programming with Large Language Models
- Murphys Laws of AI Alignment: Why the Gap Always Wins
- Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment
- Towards a Unified View of Large Language Model Post-Training
- Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
- Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology
- Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- On Aligning Prediction Models with Clinical Experiential Learning: A Prostate Cancer Case Study
- HAMSA: Hijacking Aligned Compact Models via Stealthy Automation
- On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
- NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
- MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
- SPFT-SQL: Enhancing Large Language Model for Text-to-SQL Parsing by Self-Play Fine-Tuning
- SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment
- A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models
- Measuring How (Not Just Whether) VLMs Build Common Ground
- Beyond Interpretability: Exploring the Comprehensibility of Adaptive Video Streaming through Large Language Models
- AI-in-the-Loop: Privacy Preserving Real-Time Scam Detection and Conversational Scambaiting by Leveraging LLMs and Federated Learning
- SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
- PediatricsMQA: a Multi-modal Pediatrics Question Answering Benchmark
- Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
- Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
- ChatGPT-generated texts show authorship traits that identify them as non-human
- TraceLLM: Security Diagnosis Through Traces and Smart Contracts in Ethereum
- Advancing SLM Tool-Use Capability using Reinforcement Learning
- Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
- Towards Reasoning for PDE Foundation Models: A Reward-Model-Driven Inference-Time-Scaling Algorithm
- Scaling behavior of large language models in emotional safety classification across sizes and tasks
- Omnidirectional Spatial Modeling from Correlated Panoramas
- Generative KI für TA
- The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
- Re-evaluating LLM-based Heuristic Search: A Case Study on the 3D Packing Problem
- Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving
- DCPO: Dynamic Clipping Policy Optimization
- GRAM-R2: Self-Training Generative Foundation Reward Models for Reward Reasoning
- ChatOps for microservice systems: A low-code approach using service composition and large language models
- Abex-rat: Synergizing Abstractive Augmentation and Adversarial Training for Classification of Occupational Accident Reports
- Relative Trajectory Balance is equivalent to Trust-PCL
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency
- Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic
- Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
- Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation
- Generative Goal Modeling
- On the Alignment of Large Language Models with Global Human Opinion
- Reinforcement Learning for Machine Learning Engineering Agents
- CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
- QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
- Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
- Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
- The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
- Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMs
- MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
- Confident, Calibrated, or Complicit: Probing the Trade-offs between Safety Alignment and Ideological Bias in Language Models in Detecting Hate Speech
- Political Ideology Shifts in Large Language Models
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
- Neural Models and Language Model Prompting for the Multidimensional Evaluation of Open-Ended Conversations
- RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
- LLM-Assisted Iterative Evolution with Swarm Intelligence Toward SuperBrain
- Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
- Open Data Synthesis For Deep Research
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- SABR: A Stable Adaptive Bitrate Framework Using Behavior Cloning Pretraining and Reinforcement Learning Fine-Tuning
- GIER: Gap-Driven Self-Refinement for Large Language Models
- VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
- Modeling Motivated Reasoning in Law: Evaluating Strategic Role Conditioning in LLM Summarization
- Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
- Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
- SHERPA: A Model-Driven Framework for Large Language Model Execution
- Learn from What We HAVE: History-Aware VErifier that Reasons about Past Interactions Online
- Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
- Benchmarking GPT-5 in Radiation Oncology: Measurable Gains, but Persistent Need for Expert Oversight
- Introduction to the Analysis of Probabilistic Decision-Making Algorithms
- Challenges and Applications of Large Language Models: A Comparison of GPT and DeepSeek family of models
- Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards
- SoK: Exposing the Generation and Detection Gaps in LLM-Generated Phishing
- Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
- MultiFluxAI Enhancing Platform Engineering with Advanced Agent-Orchestrated Retrieval Systems
- Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- Personality Matters: User Traits Predict LLM Preferences in Multi-Turn Collaborative Tasks
- BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions
- InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity
- AI Reasoning Models for Problem Solving in Physics
- Language-Enhanced Mobile Manipulation for Efficient Object Search in Indoor Environments
- Bridging Minds and Machines: Toward an Integration of AI and Cognitive Science
- NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
- Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs
- TCIA: A Task-Centric Instruction Augmentation Method for Instruction Finetuning
- Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution
- Learning to Generate Unit Test via Adversarial Reinforcement Learning
- Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection
- Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant
- Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
- JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
- On the possibility of deep alignment
- SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models
- 11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
- Model Science: getting serious about verification, explanation and control of AI systems
- IntentionReasoner: Facilitating Adaptive LLM Safeguards through Intent Reasoning and Selective Query Refinement
- Evaluating Language Model Reasoning about Confidential Information
- CapTune: Adapting Non-Speech Captions With Anchored Generative Models
- HEAL: A Hypothesis-Based Preference-Aware Analysis Framework
- PSO-Merging: Merging Models Based on Particle Swarm Optimization
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation?
- Continuously Steering LLMs Sensitivity to Contextual Knowledge with Proxy Models
- Safety Alignment Should Be Made More Than Just A Few Attention Heads
- Position: The Pitfalls of Over-Alignment: Overly Caution Health-Related Responses From LLMs are Unethical and Dangerous
- Skill-based Explanations for Serendipitous Course Recommendation
- Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities
- Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
- MotionFlux: Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment
- Learning Game-Playing Agents with Generative Code Optimization
- Towards 6G Intelligence: The Role of Generative AI in Future Wireless Networks
- Do MLLMs Really Understand the Charts?
- Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
- Ensemble Debates with Local Large Language Models for AI Alignment
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- Active Query Selection for Crowd-Based Reinforcement Learning
- From Bits to Boardrooms: A Cutting-Edge Multi-Agent LLM Framework for Business Excellence
- HAEPO: History-Aggregated Exploratory Policy Optimization
- Recycling History: Efficient Recommendations from Contextual Dueling Bandits
- CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
- Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement
- Beyond Quality: Unlocking Diversity in Ad Headline Generation with Large Language Models
- Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks
- RLMR: Reinforcement Learning with Mixed Rewards for Creative Writing
- LaQual: A Novel Framework for Automated Evaluation of LLM App Quality
- TrackRec: Iterative Alternating Feedback with Chain-of-Thought via Preference Alignment for Recommendation
- Better Language Model-Based Judging Reward Modeling through Scaling Comprehension Boundaries
- Stand on The Shoulders of Giants: Building JailExpert from Previous Attack Experience
- Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
- SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis
- SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
- SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
- Transduction is All You Need for Structured Data Workflows
- Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- Open-Universe Assistance Games
- Mapping the Course for Prompt-based Structured Prediction
- Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
- Reinforcement learning entangling operations on spin qubits
- Linear Preference Optimization: Decoupled Gradient Control via Absolute Regularization
- Distribution-Guided Auto-Encoder for User Multimodal Interest Cross Fusion
- In2x at WMT25 Translation Task
- DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization
- Automated Optimization Modeling through Expert-Guided Large Language Model Reasoning
- DEPTH: Hallucination-Free Relation Extraction via Dependency-Aware Sentence Simplification and Two-tiered Hierarchical Refinement
- NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
- BioLORD-2023: semantic textual representations fusing large language models and clinical knowledge graph insights
- Your Reward Function for RL is Your Best PRM for Search: Unifying RL and Search-Based TTS
- Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper
- Embedding Democratic Values into Social Media AIs via Societal Objective Functions
- Incident Analysis for AI Agents
- Learning from Preferences and Mixed Demonstrations in General Settings
- ChronoLLM: Customizing Language Models for Physics-Based Simulation Code Generation
- AI Testing Should Account for Sophisticated Strategic Behaviour
- LLMind 2.0: Distributed IoT Automation with Natural Language M2M Communication and Lightweight LLM Agents
- CIA+TA Risk Assessment for AI Reasoning Vulnerabilities
- Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
- Expertise-aware Multi-LLM Recruitment and Collaboration for Medical Decision-Making
- Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
- LM Agents May Fail to Act on Their Own Risk Knowledge
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- MAVIS: Multi-Objective Alignment via Inference-Time Value-Guided Selection
- ALIGN: Word Association Learning for Cultural Alignment in Large Language Models
- LENS: Learning to Segment Anything with Unified Reinforced Reasoning
- CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection
- One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning
- DPad: Efficient Diffusion Language Models with Suffix Dropout
- Graph Concept Bottleneck Models
- FLAIR: Feedback Learning for Adaptive Information Retrieval
- Stands to Reason: Investigating the Effect of Reasoning on Idiomaticity Detection
- AI Agents for Photonic Integrated Circuit Design Automation
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Hallucinations in medical devices
- Involuntary Jailbreak: On Self-Prompting Attacks
- Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning
- Consiglieres in the Shadow: Understanding the Use of Uncensored Large Language Models in Cybercrimes
- RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models
- Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
- MDPO: Overcoming the Training-Inference Divide of Masked Diffusion Language Models
- Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback
- Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position
- The Cultural Gene of Large Language Models: A Study on the Impact of Cross-Corpus Training on Model Values and Biases
- Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions
- Mitigating Jailbreaks with Intent-Aware LLMs
- SafeCtrl: Region-Based Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress
- QuarkMed Medical Foundation Model Technical Report
- In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
- Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings
- Learning Wisdom from Errors: Promoting LLM's Continual Relation Learning through Exploiting Error Cases
- Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation
- Controlling Multimodal LLMs via Reward-guided Decoding
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- AI Agentic Programming: A Survey of Techniques, Challenges, and Opportunities
- Preference Models assume Proportional Hazards of Utilities
- Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions
- From Clicks to Preference: A Multi-stage Alignment Framework for Generative Query Suggestion in Conversational System
- Feedback Indicators: The Alignment between Llama and a Teacher in Language Learning
- Fusing Rewards and Preferences in Reinforcement Learning
- HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
- Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
- Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
- FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
- Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
- Hard Examples Are All You Need: Maximizing GRPO Post-Training Under Annotation Budgets
- DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
- SafeConstellations: Mitigating Over-Refusals in LLMs Through Task-Aware Representation Steering
- Tailoring Diagnostic Modeling to Individual Learners: Personalized Distractor Generation via MCTS-Guided Reasoning Reconstruction
- Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
- Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes
- Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends
- CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
- Reinforced Language Models for Sequential Decision Making
- Agentic Design Review System
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Large Model Empowered Embodied AI: A Survey on Decision-Making and Embodied Learning
- Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
- ReviewRL: Towards Automated Scientific Review with RL
- Artificial Emotion: A Survey of Theories and Debates on Realising Emotion in Artificial Intelligence
- LingVarBench: Benchmarking LLM for Automated Named Entity Recognition in Structured Synthetic Spoken Transcriptions
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- Amazon Nova AI Challenge -- Trusted AI: Advancing secure, AI-assisted software development
- Perturbed Public Voices (P2V): A Dataset for Robust Audio Deepfake Detection
- Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education
- BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
- What are the limits to biomedical research acceleration through general-purpose AI?
- MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
- Slow Tuning and Low-Entropy Masking for Safe Chain-of-Thought Distillation
- On Negative-aware Preference Optimization for Recommendation
- The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
- User-centric Subjective Leaderboard by Customizable Reward Modeling
- COMPEER: Controllable Empathetic Reinforcement Reasoning for Emotional Support Conversation
- ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
- Affordances of Sketched Notations for Multimodal UI Design and Development Tools
- Reducing Cognitive Overhead in Tool Use via Multi-Small-Agent Reinforcement Learning
- BiasGym: A Simple and Generalizable Framework for Analyzing and Removing Biases through Elicitation
- Interpretable Reward Model via Sparse Autoencoder
- STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
- A Survey on Training-free Alignment of Large Language Models
- Special-Character Adversarial Attacks on Open-Source Language Model
- Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
- DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
- Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
- Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
- From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
- PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
- Vision Generalist Model: A Survey
- ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation
- DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
- Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
- Generating Query-Relevant Document Summaries via Reinforcement Learning
- SAEMark: Steering Personalized Multilingual LLM Watermarks with Sparse Autoencoders
- Reinforcement Learning for Large Model: A Survey
- WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
- \(X\)-evolve: Solution space evolution powered by large language models
- Learning to Align, Aligning to Learn: A Unified Approach for Self-Optimized Alignment
- ThinkTuning: Instilling Cognitive Reflections without Distillation
- From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR
- Careful Queries, Credible Results: Teaching RAG Models Advanced Web Search Tools with Reinforcement Learning
- Large Language Models for Subjective Language Understanding: A Survey
- Grid2Guide: A* Enabled Small Language Model for Indoor Navigation
- Vision-Based Localization and LLM-based Navigation for Indoor Environments
- Data Selection for LLM Alignment Using Fine-Grained Preferences
- Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints
- NeuroDx-LM: A Clinical Large-Scale Model for EEG-based Neurological Disorder Detection
- Pareto Multi-Objective Alignment for Language Models
- Invert4TVG: A Temporal Video Grounding Framework with Inversion Tasks for Enhanced Action Understanding
- Think Before You Talk: Enhancing Meaningful Dialogue Generation in Full-Duplex Speech Language Models with Planning-Inspired Text Guidance
- Improved Personalized Headline Generation via Denoising Fake Interests from Implicit Feedback
- A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection
- A Principled Loss Function for Direct Language Model Alignment
- Pref-GUIDE: Continual Policy Learning from Real-Time Human Feedback via Preference-Based Learning
- AMFT: Aligning LLM Reasoners by Meta-Learning the Optimal Imitation-Exploration Balance
- AR-GRPO: Training Autoregressive Image Generation Models via Reinforcement Learning
- Highlight All the Phrases: Enhancing LLM Transparency through Visual Factuality Indicators
- Many-Turn Jailbreaking
- Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models
- The NordDRG AI Benchmark for Large Language Models
- Bridging Cultural Nuances in Dialogue Agents through Cultural Value Surveys
- Demystifying Feature Requests: Leveraging LLMs to Refine Feature Requests in Open-Source Software
- Enhancing second language speaking assessment: Integrating large language models for Finnish and Finland Swedish proficiency scoring
- HapticLLaMA: A Multimodal Sensory Language Model for Haptic Captioning
- The Fair Game: Auditing & Debiasing AI Algorithms Over Time
- Sample-efficient LLM Optimization with Reset Replay
- LLM Unlearning Without an Expert Curated Dataset
- Towards Integrated Alignment
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generation
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
- Position: Intelligent Coding Systems Should Write Programs with Justifications
- A Framework for Inherently Safer AGI through Language-Mediated Active Inference
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
- Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models
- Iterative Learning of Computable Phenotypes for Treatment Resistant Hypertension using Large Language Models
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction
- Mixed-Initiative Dialog for Human-Robot Collaborative Manipulation
- The World According to LLMs: How Geographic Origin Influences LLMs' Entity Deduction Capabilities
- Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
- Can Language Models Critique Themselves? Investigating Self-Feedback for Retrieval Augmented Generation at BioASQ 2025
- mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering
- Decision-Making with Deliberation: Meta-reviewing as a Document-grounded Dialogue
- ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
- QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
- AI-assisted JSON Schema Creation and Mapping
- Posterior-GRPO: Rewarding Reasoning Processes in Code Generation
- Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
- SPaRFT: Self-Paced Reinforcement Fine-Tuning for Large Language Models
- Can Large Language Models Integrate Spatial Data? Empirical Insights into Reasoning Strengths and Computational Weaknesses
- Decoupling Understanding from Reasoning via Problem Space Mapping for Small-Scale Model Reasoning
- Echo: Decoupling Inference and Training for Large-Scale RL Alignment on Heterogeneous Swarms
- Root Cause Analysis Training for Healthcare Professionals With AI-Powered Virtual Simulation: A Proof-of-Concept
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- FaST: Feature-aware Sampling and Tuning for Personalized Preference Alignment with Limited Data
- Multi-module GRPO: Composing Policy Gradients and Prompt Optimization for Language Model Programs
- IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
- P-Aligner: Enabling Pre-Alignment of Language Models via Principled Instruction Synthesis
- TRAIL: Joint Inference and Refinement of Knowledge Graphs with Large Language Models
- SimInstruct: A Responsible Tool for Collecting Scaffolding Dialogues Between Experts and LLM-Simulated Novices
- Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
- Method-Based Reasoning for Large Language Models: Extraction, Reuse, and Continuous Improvement
- Large Language Model's Multi-Capability Alignment in Biomedical Domain
- PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
- LoRA is All You Need for Safety Alignment of Reasoning LLMs
- KG-Augmented Executable CoT for Mathematical Coding
- AttriLens-Mol: Attribute Guided Reinforcement Learning for Molecular Property Prediction with Large Language Models
- Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
- GeoSR: Cognitive-Agentic Framework for Probing Geospatial Knowledge Boundaries via Iterative Self-Refinement
- Generative Bid Shading in Real-Time Bidding Advertising
- Controllable Hybrid Captioner for Improved Long-form Video Understanding
- Sotopia-RL: Reward Design for Social Intelligence
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- DiWA: Diffusion Policy Adaptation with World Models
- LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations at eBay
- LaTCoder: Converting Webpage Design to Code with Layout-as-Thought
- MultiRAG: A Knowledge-guided Framework for Mitigating Hallucination in Multi-source Retrieval Augmented Generation
- Hide and Seek with LLMs: An Adversarial Game for Sneaky Error Generation and Self-Improving Diagnosis
- EvaDrive: Evolutionary Adversarial Policy Optimization for End-to-End Autonomous Driving
- Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling
- V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion Models
- GTPO: Stabilizing Group Relative Policy Optimization via Gradient and Entropy Control
- Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
- Fine-Tuning Text-to-Speech Diffusion Models Using Reinforcement Learning with Human Feedback
- Token-Level Precise Attack on RAG: Searching for the Best Alternatives to Mislead Generation
- ORVIT: Near-Optimal Online Distributionally Robust Reinforcement Learning
- Survey of Large Language Models in Extended Reality: Technical Paradigms and Application Frontiers
- Machine culture
- On the Evaluation of Large Language Models in Multilingual Vulnerability Repair
- PLoRA: Efficient LoRA Hyperparameter Tuning for Large Models
- CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal Knowledge
- An Efficient and Adaptive Next Edit Suggestion Framework with Zero Human Instructions in IDEs
- CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
- Traffic-R1: Reinforced LLMs Bring Human-Like Reasoning to Traffic Signal Control Systems
- CAAD: Context-Aware Adaptive Decoding for Truthful Text Generation
- A Survey on AgentOps: Categorization, Challenges, and Future Directions
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
- MArgE: Meshing Argumentative Evidence from Multiple Large Language Models for Justifiable Claim Verification
- Uni-Layout: Integrating Human Feedback in Unified Layout Generation and Evaluation
- SAMPO-Path: Segmentation Intent-Aligned Preference Optimization for Pathology Foundation Model Segmentation
- TIBSTC-CoT: A Multi-Domain Instruction Dataset for Chain-of-Thought Reasoning in Language Models
- Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
- A Survey on Data Security in Large Language Models
- MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMs
- Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
- Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
- Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention
- Censored Sampling for Topology Design: Guiding Diffusion with Human Preferences
- A Theory of Adaptive Scaffolding for LLM-Based Pedagogical Agents
- TeSent: A Benchmark Dataset for Fairness-aware Explainable Sentiment Classification in Telugu
- From Query to Logic: Ontology-Driven Multi-Hop Reasoning in LLMs
- MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
- Adaptive Content Restriction for Large Language Models via Suffix Optimization
- Benchmarking and Bridging Emotion Conflicts for Multimodal Emotion Reasoning
- RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models
- Bias Association Discovery Framework for Open-Ended LLM Generations
- Provably Secure Retrieval-Augmented Generation
- A Note on Code Quality Score: LLMs for Maintainable Large Codebases
- Better Call Claude: Can LLMs Detect Changes of Writing Style?
- Activation-Guided Local Editing for Jailbreaking Attacks
- Foundations of Interpretable Models
- ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety
- Thinking Machines: Mathematical Reasoning in the Age of LLMs
- When Relevance Meets Novelty: Dual-Stable Periodic Optimization for Serendipitous Recommendation
- AutoDebias: Automated Framework for Debiasing Text-to-Image Models
- Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation
- Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
- MetaAgent: Toward Self-Evolving Agent via Tool Meta-Learning
- MAO-ARAG: Multi-Agent Orchestration for Adaptive Retrieval-Augmented Generation
- Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report
- Lucy: edgerunning agentic web search on mobile with machine generated task vectors
- Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
- Calibrated Language Models and How to Find Them with Label Smoothing
- ITDR: An Instruction Tuning Dataset for Enhancing Large Language Models in Recommendations
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
- The Other Mind: How Language Models Exhibit Human Temporal Cognition
- GraphRAG-R1: Graph Retrieval-Augmented Generation with Process-Constrained Reinforcement Learning
- Causal2Vec: Improving Decoder-only LLMs as Versatile Embedding Models
- Text-to-SQL Task-oriented Dialogue Ontology Construction
- What's Taboo for You? - An Empirical Evaluation of LLMs Behavior Toward Sensitive Content
- AutoBridge: Automating Smart Device Integration with Centralized Platform
- BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
- How Far Are AI Scientists from Changing the World?
- T-Detect: Tail-Aware Statistical Normalization for Robust Detection of Adversarial Machine-Generated Text
- MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
- DynaSwarm: Dynamically Graph Structure Selection for LLM-based Multi-agent System
- Scalable Multi-Task Reinforcement Learning for Generalizable Spatial Intelligence in Visuomotor Agents
- Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR
- Rethinking Evidence Hierarchies in Medical Language Benchmarks: A Critical Evaluation of HealthBench
- Open-Source Agentic Hybrid RAG Framework for Scientific Literature Review
- Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
- Opportunities and Challenges of LLMs in Education: An NLP Perspective
- OFCnetLLM: Large Language Model for Network Monitoring and Alertness
- SIA: Enhancing Safety via Intent Awareness for Vision-Language Models
- ShortFT: Diffusion Model Alignment via Shortcut-based Fine-Tuning
- Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
- Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems
- Learning to Extract Rational Evidence via Reinforcement Learning for Retrieval-Augmented Generation
- Heartificial Intelligence: Exploring Empathy in Language Models
- Strategic Deflection: Defending LLMs from Logit Manipulation
- Improving Generative Ad Text on Facebook using Reinforcement Learning
- Reasoning Language Models for Root Cause Analysis in 5G Wireless Networks
- Multimodal Video Emotion Recognition with Reliable Reasoning Priors
- Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
- Introducing HALC: A general pipeline for finding optimal prompting strategies for automated coding with LLMs in the computational social sciences
- MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
- AgriEval: A Comprehensive Chinese Agricultural Benchmark for Large Language Models
- Libra: Assessing and Improving Reward Model by Learning to Think
- Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models
- Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
- Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs
- Towards Locally Deployable Fine-Tuned Causal Large Language Models for Mode Choice Behaviour
- GovRelBench:A Benchmark for Government Domain Relevance
- TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
- Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
- Flow Matching Policy Gradients
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
- AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
- Uncovering Gradient Inversion Risks in Practical Language Model Training
- Kimi K2: Open Agentic Intelligence
- Metaphor and Large Language Models: When Surface Features Matter More than Deep Understanding
- CodeNER: Code Prompting for Named Entity Recognition
- Length Representations in Large Language Models
- Cultivating Helpful, Personalized, and Creative AI Tutors: A Framework for Pedagogical Alignment using Reinforcement Learning
- The Blessing and Curse of Dimensionality in Safety Alignment
- A Novel Self-Evolution Framework for Large Language Models
- ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling
- IM-Chat: A Multi-agent LLM Framework Integrating Tool-Calling and Diffusion Modeling for Knowledge Transfer in Injection Molding Industry
- Post-Completion Learning for Language Models
- NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding
- SGPO: Self-Generated Preference Optimization based on Self-Improver
- LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
- Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality
- The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models
- MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
- Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering
- SDD: Self-Degraded Defense against Malicious Fine-tuning
- Can Language Models Discover Scaling Laws?
- Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
- PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
- Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory
- KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
- The Impact of Language Mixing on Bilingual LLM Reasoning
- Leveraging Fine-Tuned Large Language Models for Interpretable Pancreatic Cystic Lesion Feature Extraction and Risk Categorization
- Ultracoarse Equilibria and Ordinal-Folding Dynamics in Operator-Algebraic Models of Infinite Multi-Agent Games
- Diverse LLMs or Diverse Question Interpretations? That is the Ensembling Question
- Integrating LLM in Agent-Based Social Simulation: Opportunities and Challenges
- AGORA: Incentivizing Group Emergence Capability in LLMs via Group Distillation
- Debating Truth: Debate-driven Claim Verification with Multiple Large Language Model Agents
- Legal Document Summarization: Enhancing Judicial Efficiency through Automation Detection
- Adaptive Cluster Collaborativeness Boosts LLMs Medical Decision Support Capacity
- Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
- Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
- PurpCode: Reasoning for Safer Code Generation
- Adaptive Learning Systems: Personalized Curriculum Design Using LLM-Powered Analytics
- Jailbreaking Large Language Diffusion Models: Revealing Hidden Safety Flaws in Diffusion-Based Text Generation
- DxHF: Providing High-Quality Human Feedback for LLM Alignment via Interactive Decomposition
- TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
- How Well Do LLMs Predict Prerequisite Skills? Zero-Shot Comparison to Expert-Defined Concepts
- FinDPO: Financial Sentiment Analysis for Algorithmic Trading through Preference Optimization of LLMs
- Revisiting LLM Reasoning via Information Bottleneck
- Towards Effective Human-in-the-Loop Assistive AI Agents
- Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
- Technical Report of TeleChat2, TeleChat2.5 and T1
- StyleAdaptedLM: Enhancing Instruction Following Models with Efficient Stylistic Transfer
- Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
- Exploring the Impact of Instruction-Tuning on LLM's Susceptibility to Misinformation
- AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI
- Evaluation of Coding Schemes for Transformer-based Gene Sequence Modeling
- WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization
- E.A.R.T.H.: Structuring Creative Evolution through Model Error in Generative AI
- Agent WARPP: Workflow Adherence via Runtime Parallel Personalization
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
- LEKIA: Expert-Aligned AI Behavior Design for High-Risk Human-AI Interactions
- Structured Chain-of-Thought Prompting for Code Generation
- AgentFly: Extensible and Scalable Reinforcement Learning for LM Agents
- Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
- Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization
- MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
- Acoda: Adversarial Code Obfuscation for Defending against LLM-based Analysis
- Theoretical Foundations and Mitigation of Hallucination in Large Language Models
- Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs
- CodegenBench: Can LLMs Write Efficient Code Across Architectures?
- A Primer in Post-Training Reasoning Data: What We Know About How It Works
- A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
- Omissive Bias in Religious Representation: Benchmarking LLM Answers to Everyday Ethical Decision-making
- QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks
- MeMo: Memory as a Model
- Image Generators are Generalist Vision Learners
- MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
- BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLM
- PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
- VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning
- Generative Distribution Distillation
- DPBench: Structural Determinants of Multi-Agent LLM Coordination Under Simultaneous Resource Contention
- Shaping capabilities with token-level data filtering
- Statistical and Algorithmic Foundations of Reinforcement Learning
- A Hybrid Framework for Subject Analysis: Integrating Embedding-Based Regression Models with Large Language Models
- Characterizing Communication Patterns in Distributed Large Language Model Inference
- The Levers of Political Persuasion with Conversational AI
- Critique of impure reason: Unveiling the reasoning behaviour of medical large language models
- URPO: A Unified Reward & Policy Optimization Framework for Large Language Models
- The Geometry of Harmfulness in LLMs through Subconcept Probing
- CLARIFID: Improving Radiology Report Generation by Reinforcing Clinically Accurate Impressions and Enforcing Detailed Findings
- On a few pitfalls in KL divergence gradient estimation for RL
- Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning
- R4ec: A Reasoning, Reflection, and Refinement Framework for Recommendation Systems
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models
- Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?
- Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
- The Ever-Evolving Science Exam
- Multi-Agent Reinforcement Learning for Sample-Efficient Deep Neural Network Mapping
- WakenLLM: Evaluating Reasoning Potential and Stability in LLMs via Fine-Grained Benchmarking
- Playing repeated games with large language models
- A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
- The Judge Variable: Challenging Judge-Agnostic Legal Judgment Prediction
- Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
- RePO: Replay-Enhanced Policy Optimization
- A Comprehensive Review on Harnessing Large Language Models to Overcome Recommender System Challenges
- PrefPalette: Personalized Preference Modeling with Latent Attributes
- Automating Steering for Safe Multimodal Large Language Models
- Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
- Model-free Reinforcement Learning for Model-based Control: Towards Safe, Interpretable and Sample-efficient Agents
- DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
- Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
- AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
- Learning to summarize user information for personalized reinforcement learning from human feedback
- The Serial Scaling Hypothesis
- Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
- Granular feedback merits sophisticated aggregation
- Advancing Conversational Diagnostic AI with Multimodal Reasoning
- LLMs Encode Harmfulness and Refusal Separately
- QuRe: Query-Relevant Retrieval through Hard Negative Sampling in Composed Image Retrieval
- Xiangqi-R1: Enhancing Spatial Strategic Reasoning in LLMs for Chinese Chess via Reinforcement Learning
- Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
- TFGIN: Tight-Fitting Graph Inference Network for Table-based Fact Verification
- Agentic Neural Networks: Self-Evolving Multi-Agent Systems via Textual Backpropagation
- Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning
- Language
- Learning to Reason Across Parallel Samples for LLM Reasoning
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering
- AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin
- Olica: Efficient Structured Pruning of Large Language Models without Retraining
- Product vs. Process: Exploring EFL Students' Editing of AI-Generated Text for Expository Writing
- ORFS-agent: Tool-Using Agents for Chip Design Optimization
- ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
- ThinkQE: Query Expansion via an Evolving Thinking Process
- FZOO: Fast Zeroth-Order Optimizer for Fine-Tuning Large Language Models towards Adam-Scale Speed
- Mitigating Reward Over-optimization in Direct Alignment Algorithms with Importance Sampling
- Internal Value Alignment in Large Language Models through Controlled Value Vector Activation
- Enhancing Safe and Controllable Protein Generation via Knowledge Preference Optimization
- Aligned Query Expansion: Efficient Query Expansion for Information Retrieval through LLM Alignment
- A Survey on Large Language Models for Mathematical Reasoning
- Auto-Formulating Dynamic Programming Problems with Large Language Models
- Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
- SystolicAttention: Fusing FlashAttention within a Single Systolic Array
- Towards Practical Benchmarking of Data Cleaning Techniques: On Generating Authentic Errors via Large Language Models
- How Many Instructions Can LLMs Follow at Once?
- Multi-Armed Sampling Problem and the End of Exploration
- Intra-Trajectory Consistency for Reward Modeling
- CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
- PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
- HKGAI-V1: Towards Regional Sovereign Large Language Model for Hong Kong
- Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
- Aligning Generative Speech Enhancement with Human Preferences via Direct Preference Optimization
- A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
- Factors affecting the in-context learning abilities of LLMs for dialogue state tracking
- What Should Feature Distillation Transfer in LLMs? A Task-Tangent Geometry View
- Retention analysis of edited knowledge after fine-tuning
- DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models
- Foundation Model Driven Robotics: A Comprehensive Review
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Turning the Tide: Repository-based Code Reflection
- Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models
- Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
- Can AI Rely on the Systematicity of Truth? The Challenge of Modelling Normative Domains
- Intention-Conditioned Flow Occupancy Models
- ViSP: A PPO-Driven Framework for Sarcasm Generation with Contrastive Learning
- RedOne: Revealing Domain-specific LLM Post-Training in Social Networking Services
- Fine-tuning Large Language Model for Automated Algorithm Design
- GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO
- Compute Requirements for Algorithmic Innovation in Frontier AI Models
- DATE-LM: Benchmarking Data Attribution Evaluation for Large Language Models
- Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models
- Adversarial Activation Patching: A Framework for Detecting and Mitigating Emergent Deception in Safety-Aligned Transformers
- Detecting and Pruning Prominent but Detrimental Neurons in Large Language Models
- Self-Improving Model Steering
- Lumos-1: On Autoregressive Video Generation from a Unified Model Perspective
- NeuralOS: Towards Simulating Operating Systems via Neural Generative Models
- One Token to Fool LLM-as-a-Judge
- Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
- From Language to Logic: A Bi-Level Framework for Structured Reasoning
- Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework
- A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities
- Inference-Time Scaling of Diffusion Language Models with Particle Gibbs Sampling
- Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
- SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven Manner
- CMER: A Context-Aware Approach for Mining Ethical Concern-related App Reviews
- Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
- A Dynamic Stackelberg Game Framework for Agentic AI Defense Against LLM Jailbreaking
- CTRLS: Chain-of-Thought Reasoning via Latent State-Transition
- Audit, Alignment, and Optimization of LM-Powered Subroutines with Application to Public Comment Processing
- Low-rank Momentum Factorization for Memory Efficient Training
- Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
- Explore Beyond the Boundary Using Entropic Information
- Studying quantization trade-offs for efficient inference deployment in machine translation
- PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
- Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing
- Scaling RL to Long Videos
- Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
- GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
- Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
- Why is Your Language Model a Poor Implicit Reward Model?
- Divergence Minimization Preference Optimization for Diffusion Model Alignment
- 5C Prompt Contracts: A Minimalist, Creative-Friendly, Token-Efficient Design Framework for Individual and SME LLM Usage
- Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report
- Winning and losing with Artificial Intelligence: What public discourse about ChatGPT tells us about how societies make sense of technological change
- TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text
- From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization
- A Mathematical Theory of Discursive Networks
- Large Language Model for Extracting Complex Contract Information in Industrial Scenes
- Robust Multimodal Large Language Models Against Modality Conflict
- First Return, Entropy-Eliciting Explore
- Could the Road to Grounded, Neuro-symbolic AI be Paved with Words-as-Classifiers?
- CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
- Evolution without Large Models: Training Language Model with Task Principles
- When Transformers Meet Recommenders: Integrating Self-Attentive Sequential Recommendation with Fine-Tuned LLMs
- TuneShield: Mitigating Toxicity in Conversational AI while Fine-tuning on Untrusted Data
- How Not to Detect Prompt Injections with an LLM
- Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
- Detecting Proxy Gaming in RL and LLM Alignment via Evaluator Stress Tests
- The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
- Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering
- The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
- OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation
- Affective-ROPTester: Capability and Bias Analysis of LLMs in Predicting Retinopathy of Prematurity
- TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
- Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
- Data Compressibility Quantifies LLM Memorization
- MolFORM: Multi-modal Flow Matching for Structure-Based Drug Design
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Logit Reweighting for Topic-Focused Summarization
- Steering Information Utility in Key-Value Memory for Language Model Post-Training
- An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
- Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
- LLM-based Question-Answer Framework for Sensor-driven HVAC System Interaction
- Why We Feel What We Feel: Joint Detection of Emotions and Their Opinion Triggers in E-commerce
- Interpretable Reward Modeling with Active Concept Bottlenecks
- A Query-Aware Multi-Path Knowledge Graph Fusion Approach for Enhancing Retrieval-Augmented Generation in Large Language Models
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
- SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?
- Who's the Mole? Modeling and Detecting Intention-Hiding Malicious Agents in LLM-Based Multi-Agent Systems
- Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message
- wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
- TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter
- Discrete Diffusion Trajectory Alignment via Stepwise Decomposition
- Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models
- Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
- Pre-Trained Policy Discriminators are General Reward Models
- Unveiling the Potential of Diffusion Large Language Model in Controllable Generation
- GradOT: Training-free Gradient-preserving Offsite-tuning for Large Language Models
- Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning
- Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
- Fairness Evaluation of Large Language Models in Academic Library Reference Services
- ESSA: Evolutionary Strategies for Scalable Alignment
- HLStrans: Dataset for C-to-HLS Hardware Code Synthesis
- A Technical Survey of Reinforcement Learning Techniques for Large Language Models
- Easy Dataset: A Unified and Extensible Framework for Synthesizing LLM Fine-Tuning Data from Unstructured Documents
- Losing our Tail, Again: (Un)Natural Selection & Multilingual LLMs
- Enhancing Adaptive Behavioral Interventions with LLM Inference from Participant-Described States
- Rethinking and Exploring String-Based Malware Family Classification in the Era of LLMs and RAG
- Economic Evaluation of LLMs
- Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
- Large Language Models for Combinatorial Optimization: A Systematic Review
- WETBench: A Benchmark for Detecting Task-Specific Machine-Generated Text on Wikipedia
- Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky
- FlexPath: Adapting Learned Connectivity Guidance to Path Preferences
- Relationship Between Trust in the AI Creator and Trust in AI Systems: The Crucial Role of AI Alignment and Steerability
- TACOS: Open Tagging and Comparative Scoring for Instruction Fine-Tuning Data Selection
- MemOS: A Memory OS for AI System
- Understanding Knowledge Transferability for Transfer Learning: A Survey
- Assessing Small Language Models for Code Generation: An Empirical Study with Benchmarks
- How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models
- Requirements Elicitation Follow-Up Question Generation
- ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
- PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
- MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion
- Improving Consistency in Vehicle Trajectory Prediction Through Preference Optimization
- MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
- Listwise Preference Alignment Optimization for Tail Item Recommendation
- Autonomous Control Leveraging LLMs: An Agentic Framework for Next-Generation Industrial Automation
- Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models
- ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning
- Data Diversification Methods In Alignment Enhance Math Performance In LLMs
- Can Artificial Intelligence solve the blockchain oracle problem? Unpacking the Challenges and Possibilities
- Energy-Based Transformers are Scalable Learners and Thinkers
- MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation
- Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation
- MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining
- Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
- Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation
- OpenTable-R1: A Reinforcement Learning Augmented Tool Agent for Open-Domain Table Question Answering
- Emotionally Intelligent Task-oriented Dialogue Systems: Architecture, Representation, and Optimisation
- Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities
- Activation Reward Models for Few-Shot Model Alignment
- Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs
- Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy
- AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
- ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
- Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
- Dynamic Strategy Adaptation in Multi-Agent Environments with Large Language Models
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications
- Multi-interaction TTS toward professional recording reproduction
- Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
- SAFER: Probing Safety in Reward Models with Sparse Autoencoder
- FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
- Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
- TransLaw: A Large-Scale Dataset and Multi-Agent Benchmark Simulating Professional Translation of Hong Kong Case Law
- Reasoning as an Adaptive Defense for Safety
- From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
- Efficient Post-Training Refinement of Latent Reasoning in Large Language Models
- Linearly Decoding Refused Knowledge in Aligned Language Models
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
- Scaling Human Judgment in Community Notes with LLMs
- Auto-TA: Towards Scalable Automated Thematic Analysis (TA) via Multi-Agent Large Language Models with Reinforcement Learning
- PBa-LLM: Privacy- and Bias-aware NLP using Named-Entity Recognition (NER)
- A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents
- Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
- PageLLM: A Multi-Grained Reward Framework for Whole-Page Optimization with Large Language Models
- Epitome: Pioneering an Experimental Platform for AI-Social Science Integration
- Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective
- The role of LLMs in theory building
- “Desired behaviors”: alignment and the emergence of a machine learning ethics
- GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors
- Bingo: Boosting Efficient Reasoning of LLMs via Dynamic and Significance-based Reinforcement Learning
- QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
- TaP: A Taxonomy-Guided Framework for Automated and Scalable Preference Data Generation
- Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
- Optimizing Conversational Product Recommendation via Reinforcement Learning
- Prompting as Scientific Inquiry
- TuCo: Measuring the Contribution of Fine-Tuning to Individual Responses of LLMs
- Real-Time Progress Prediction in Reasoning Language Models
- Masked Gated Linear Unit
- Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier Gap
- AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
- Generalist Reward Models: Found Inside Large Language Models
- Treatment, evidence, imitation, and chat
- RL4CO: an Extensive Reinforcement Learning for Combinatorial Optimization Benchmark
- A Systematic Study of Compositional Syntactic Transformer Language Models
- Video Unlearning via Low-Rank Refusal Vector
- Towards Explainable Bilingual Multimodal Misinformation Detection and Localization
- Knowledge Augmented Finetuning Matters in both RAG and Agent Based Dialog Systems
- FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 Precision
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
- FairMarket-RL: LLM-Guided Fairness Shaping for Multi-Agent Reinforcement Learning in Peer-to-Peer Markets
- VERA: Variational Inference Framework for Jailbreaking Large Language Models
- Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
- Improving Large Language Models with Concept-Aware Fine-Tuning
- Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
- A Survey of Continual Reinforcement Learning
- LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
- AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
- GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles
- The Hidden Link Between RLHF and Contrastive Learning
- RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
- SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
- Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
- More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents
- Data Efficacy for Language Model Training
- APO: Enhancing Reasoning Ability of MLLMs via Asymmetric Policy Optimization
- Active Inference AI Systems for Scientific Discovery
- HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
- From Threat to Tool: Leveraging Refusal-Aware Injection Attacks for Safety Alignment
- Explicit Preference Optimization: No Need for an Implicit Reward Model
- TRIDENT: Tri-Modal Molecular Representation Learning with Taxonomic Annotations and Local Correspondence
- Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
- LLM-guided Chemical Process Optimization with a Multi-Agent Approach
- Improving Fairness of Large Language Models in Multi-document Summarization
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
- When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
- Aligning Spoken Dialogue Models from User Interactions
- Evidence-based diagnostic reasoning with multi-agent copilot for human pathology
- Learning to Skip the Middle Layers of Transformers
- Bridging Offline and Online Reinforcement Learning for LLMs
- Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models
- Model Editing as a Double-Edged Sword: Steering Agent Ethical Behavior Toward Beneficence or Harm
- Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
- Unraveling the Potential of Diffusion Models in Small Molecule Generation
- Mitigating Gambling-Like Risk-Taking Behaviors in Large Language Models: A Behavioral Economics Approach to AI Safety
- A Survey of Behavior Foundation Model: Next-Generation Whole-Body Control System of Humanoid Robots
- Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models
- ReCode: Updating Code API Knowledge with Reinforcement Learning
- Asymmetric REINFORCE for off-Policy Reinforcement Learning: Balancing positive and negative rewards
- Leveraging AI Graders for Missing Score Imputation to Achieve Accurate Ability Estimation in Constructed-Response Tests
- Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks
- Synthesis by Design: Controlled Data Generation via Structural Guidance
- Persona Features Control Emergent Misalignment
- KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality
- Can Reasoning Help Large Language Models Capture Human Annotator Disagreement?
- Reinforcement Learning via Implicit Imitation Guidance
- Design Principles for Generative AI Applications
- FEAT: A Preference Feedback Dataset through a Cost-Effective Auto-Generation and Labeling Framework for English AI Tutoring
- Hallucination Detection with Small Language Models
- Automatic Prompt Optimization for Knowledge Graph Construction: Insights from an Empirical Study
- Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
- RecLLM-R1: A Two-Stage Training Paradigm with Reinforcement Learning and Chain-of-Thought v1
- Controlled Retrieval-augmented Context Evaluation for Long-form RAG
- Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
- HiMA-Ecom: Enabling Joint Training of Hierarchical Multi-Agent E-commerce Assistants
- CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
- MiniCPM4: Ultra-Efficient LLMs on End Devices
- ConciseHint: Boosting Efficient Reasoning via Continuous Concise Hints during Generation
- Harnessing the Power of Reinforcement Learning for Language-Model-Based Information Retriever via Query-Document Co-Augmentation
- ReDit: Reward Dithering for Improved LLM Policy Optimization
- SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
- End-to-End Spoken Grammatical Error Correction
- Generalizing vision-language models to novel domains: A comprehensive survey
- Comparative Evaluation of ChatGPT and DeepSeek Across Key NLP Tasks: Strengths, Weaknesses, and Domain-Specific Performance
- A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models
- Adaptive alert prioritisation in security operations centres via learning to defer with human feedback
- End-to-End Fine-Tuning of 3D Texture Generation using Differentiable Rewards
- NSFW-Classifier Guided Prompt Sanitization for Safe Text-to-Image Generation
- Refactoring Programs Using Large Language Models with Few-Shot Examples
- Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models
- Shrinking the Generation-Verification Gap with Weak Verifiers
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- AI Through the Human Lens: Investigating Cognitive Theories in Machine Psychology
- InspireDebate: Multi-Dimensional Subjective-Objective Evaluation-Guided Reasoning and Optimization for Debating
- TIM: A Large-Scale Dataset and large Timeline Intelligence Model for Open-domain Timeline Summarization
- Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach
- Greedy Selection under Independent Increments: A Toy Model Analysis
- LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
- Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)
- RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
- Beyond Syntax: Action Semantics Learning for App Agents
- Safe Pruning LoRA: Robust Distance-Guided Pruning for Safety Alignment in Adaptation of LLMs
- AI Safety vs. AI Security: Demystifying the Distinction and Boundaries
- Distilling On-device Language Models for Robot Planning with Minimal Human Intervention
- Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
- Research quality evaluation by AI in the era of Large Language Models: Advantages, disadvantages, and systemic effects
- Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basque
- SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification
- Reward-Agnostic Prompt Optimization for Text-to-Image Diffusion Models
- The Role of Model Confidence on Bias Effects in Measured Uncertainties for Vision-Language Models
- ReasonGRM: Enhancing Generative Reward Models through Large Reasoning Models
- Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
- Can structural correspondences ground real world representational content in Large Language Models?
- PRISON: Unmasking the Criminal Potential of Large Language Models
- GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
- AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
- Probing the Robustness of Large Language Models Safety to Latent Perturbations
- Reranking-based Generation for Unbiased Perspective Summarization
- Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
- From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation
- Arch-Router: Aligning LLM Routing with Human Preferences
- StoryWriter: A Multi-Agent Framework for Long Story Generation
- Steering Your Diffusion Policy with Latent Space Reinforcement Learning
- AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
- Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
- LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
- Intelligent Assistants for the Semiconductor Failure Analysis with LLM-Based Planning Agents
- Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact
- AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need
- ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning
- video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
- Gradients: When Markets Meet Fine-tuning -- A Distributed Approach to Model Optimisation
- Representation Consistency for Accurate and Coherent LLM Answer Aggregation
- Quantile Regression with Large Language Models for Price Prediction
- Learning-Time Encoding Shapes Unlearning in LLMs
- Modeling the One-to-Many Property in Open-Domain Dialogue with LLMs
- Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts
- Lessons from Training Grounded LLMs with Verifiable Rewards
- RePCS: Diagnosing Data Memorization in LLM-Powered Retrieval-Augmented Generation
- From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
- Reward Models in Deep Reinforcement Learning: A Survey
- SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture
- SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models
- LLM Jailbreak Oracle
- MDBench: A Synthetic Multi-Document Reasoning Benchmark Generated with Knowledge Guidance
- Large Language Models -- the Future of Fundamental Physics?
- Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality
- Probabilistic Aggregation and Targeted Embedding Optimization for Collective Moral Reasoning in Large Language Models
- Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding
- Mobile Application Review Summarization using Chain of Density Prompting
- DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimization
- FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
- Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent
- Reasoning with Exploration: An Entropy Perspective
- GRAM: A Generative Foundation Reward Model for Reward Generalization
- What Makes a Good Natural Language Prompt?
- Adultification Bias in LLMs and Text-to-Image Models
- Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation Framework
- From Tool Calling to Symbolic Thinking: LLMs in a Persistent Lisp Metaprogramming Loop
- TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization
- Adaptive Accompaniment with ReaLchords
- AviationLLM: An LLM-based Knowledge System for Aviation Training
- From General Reasoning to Domain Expertise: Uncovering the Limits of Generalization in Large Language Models
- Flexible-length Text Infilling for Discrete Diffusion Models
- Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
- Value-Free Policy Optimization via Reward Partitioning
- We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
- Implicit and Explicit Research Quality Score Probabilities from ChatGPT
- Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
- Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning
- Evaluating and Improving Robustness in Large Language Models: A Survey and Future Directions
- GeometryZero: Improving Geometry Solving for LLM with Group Contrastive Policy Optimization
- AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models
- Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
- Fake it till You Make it: Reward Modeling as Discriminative Prediction
- Meta-learning how to Share Credit among Macro-Actions
- "I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning
- Calibrated Predictive Lower Bounds on Time-to-Unsafe-Sampling in LLMs
- VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
- Discrete Diffusion in Large Language and Multimodal Models: A Survey
- Document-Level Tabular Numerical Cross-Checking: A Coarse-to-Fine Approach
- Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems
- TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning
- Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
- Mind the Web: The Security of Web Use Agents
- ASMR: Augmenting Life Scenario using Large Generative Models for Robotic Action Reflection
- Human-assisted Robotic Policy Refinement via Action Preference Optimization
- Socratic RL: A Novel Framework for Efficient Knowledge Acquisition through Iterative Reflection and Viewpoint Distillation
- OneRec Technical Report
- Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
- HauntAttack: When Attack Follows Reasoning as a Shadow
- Jailbreak Transferability Emerges from Shared Representations
- Diffusion Policies for Out-of-Distribution Generalization in Offline Reinforcement Learning
- Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
- Rethinking Hate Speech Detection on Social Media: Can LLMs Replace Traditional Models?
- Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
- Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
- History-Aware Cross-Attention Reinforcement: Self-Supervised Multi Turn and Chain-of-Thought Fine-Tuning with vLLM
- Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
- Mathesis: Towards Formal Theorem Proving from Natural Languages
- Identifying and Investigating Global News Coverage of Critical Events Such as Disasters and Terrorist Attacks
- Large Language Models Enhanced by Plug and Play Syntactic Knowledge for Aspect-based Sentiment Analysis
- Can LLMs Reconcile Knowledge Conflicts in Counterfactual Reasoning
- Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
- Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
- Bridging the Digital Divide: Small Language Models as a Pathway for Physics and Photonics Education in Underdeveloped Regions
- AlphaSteer: Learning Refusal Steering with Principled Null-Space Constraint
- Exploring the Secondary Risks of Large Language Models
- CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
- Evaluating LLM-Contaminated Crowdsourcing Data Without Ground Truth
- Detection, Classification, and Mitigation of Gender Bias in Large Language Models
- Profiling News Media for Factuality and Bias Using LLMs and the Fact-Checking Methodology of Human Experts
- Advances in LLMs with Focus on Reasoning, Adaptability, Efficiency and Ethics
- Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory
- From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment
- Diffusion-Based Electrocardiography Noise Quantification via Anomaly Detection
- Eliciting Reasoning in Language Models with Cognitive Tools
- Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
- Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models
- Personalized LLM Decoding via Contrasting Personal Preference
- Towards Understanding the Cognitive Habits of Large Reasoning Models
- Curriculum-Guided Layer Scaling for Language Model Pretraining
- Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information
- TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
- Tokenized Bandit for LLM Decoding and Alignment
- Mind the XAI Gap: A Human-Centered LLM Framework for Democratizing Explainable AI
- The Behavior Gap: Evaluating Zero-shot LLM Agents in Complex Task-Oriented Dialogs
- Fed-HeLLo: Efficient Federated Foundation Model Fine-Tuning with Heterogeneous LoRA Allocation
- Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
- How large language models can reshape collective intelligence
- AnnoDPO: Protein Functional Annotation Learning with Direct Preference Optimization
- LLM-as-a-Fuzzy-Judge: Fine-Tuning Large Language Models as a Clinical Evaluation Judge with Fuzzy Logic
- GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
- SoK: Evaluating Jailbreak Guardrails for Large Language Models
- The Alignment Trap: Complexity Barriers
- EQA-RM: A Generative Embodied Reward Model with Test-time Scaling
- Spurious Rewards: Rethinking Training Signals in RLVR
- Can We Infer Confidential Properties of Training Data from LLMs?
- Time Series Forecasting as Reasoning: A Slow-Thinking Approach with Reinforced LLMs
- Table-Text Alignment: Explaining Claim Verification Against Tables in Scientific Papers
- OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
- Assessing RAG and HyDE on 1B vs. 4B-Parameter Gemma LLMs for Personal Assistants Integretion
- Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
- Don't Pay Attention
- Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
- Poutine: Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving
- Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
- Collaborative Prediction: To Join or To Disjoin Datasets
- Pareto Optimal Code Generation
- LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
- DreamCS: Geometry-Aware Text-to-3D Generation with Unpaired 3D Reward Supervision
- Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives
- TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning
- Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
- EnerBridge-DPO: Energy-Guided Protein Inverse Folding with Markov Bridges and Direct Preference Optimization
- OneSug: The Unified End-to-End Generative Framework for E-commerce Query Suggestion
- Reasoning Models Don't Always Say What They Think
- PersonaAgent: Bridging Memory and Action for Personalized LLM Agents
- Corrector Sampling in Language Models
- WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management
- Debiasing Online Preference Learning via Preference Feature Preservation
- Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models
- Preference Learning for AI Alignment: a Causal Perspective
- Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
- The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
- SPRINT: Enabling Interleaved Planning and Parallelized Execution in Reasoning Models
- To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
- A Systematic Review of Poisoning Attacks Against Large Language Models
- Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance
- Can Theoretical Physics Research Benefit from Language Agents?
- Saffron-1: Safety Inference Scaling
- Prompting Wireless Networks: Reinforced In-Context Learning for Power Control
- Benchmarking Misuse Mitigation Against Covert Adversaries
- Distillation Robustifies Unlearning
- Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
- CoMemo: LVLMs Need Image Context with Image Memory
- Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Models
- Accelerating Sparse Transformer Inference on GPU
- Being Strong Progressively! Enhancing Knowledge Distillation of Large Language Models through a Curriculum Learning Framework
- Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
- From Bias To Improved Prompts: A Case Study of Bias Mitigation of Clone Detection Models
- MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks
- Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
- UniRes: Universal Image Restoration for Complex Degradations
- Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
- A Smooth Sea Never Made a Skilled SAILOR: Robust Imitation via Learning to Search
- SECNEURON: Reliable and Flexible Abuse Control in Local LLMs via Hybrid Neuron Encryption
- Improving Low-Resource Morphological Inflection via Self-Supervised Objectives
- LLM-Guided Scenario-based GUI Testing
- PCDVQ: Enhancing Vector Quantization for Large Language Models via Polar Coordinate Decoupling
- Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
- Normative Conflicts and Shallow AI Alignment
- Detection Method for Prompt Injection by Integrating Pre-trained Model and Heuristic Feature Engineering
- Truly Self-Improving Agents Require Intrinsic Metacognitive Learning
- DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning
- TreeRPO: Tree Relative Policy Optimization
- SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
- Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
- Search Arena: Analyzing Search-Augmented LLMs
- SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
- RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation
- ConECT Dataset: Overcoming Data Scarcity in Context-Aware E-Commerce MT
- Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
- GOLFer: Smaller LM-Generated Documents Hallucination Filter & Combiner for Query Expansion in Information Retrieval
- On the Fundamental Impossibility of Hallucination Control in Large Language Models
- Training a Scientific Reasoning Model for Chemistry
- SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models
- Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers
- Fight Fire with Fire: Defending Against Malicious RL Fine-Tuning via Reward Neutralization
- RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming
- To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
- A Large Language Model for Feasible and Diverse Population Synthesis
- Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
- Steerable Scene Generation with Post Training and Inference-Time Search
- CAD-Llama: Leveraging Large Language Models for Computer-Aided Design Parametric 3D Model Generation
- ABKD: Pursuing a Proper Allocation of the Probability Mass in Knowledge Distillation via α-β-Divergence
- Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
- Rethinking LLM Advancement: Compute-Dependent and Independent Paths to Progress
- Build Agent Advocates, Not Platform Agents
- Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving
- EchoInk-R1: Exploring Audio-Visual Reasoning in Multimodal LLMs via Reinforcement Learning
- PosterO: Structuring Layout Trees to Enable Language Models in Generalized Content-Aware Layout Generation
- Exchange of Perspective Prompting Enhances Reasoning in Large Language Models
- BPO: Revisiting Preference Modeling in Direct Preference Optimization
- Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
- Multimodal Tabular Reasoning with Privileged Structured Information
- Leveraging Reward Models for Guiding Code Review Comment Generation
- Robust Preference Optimization via Dynamic Target Margins
- Misalignment or misuse? The AGI alignment tradeoff
- Aligning Large Language Models with Implicit Preferences from User-Generated Content
- RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
- PPO in the Fisher-Rao geometry
- EpiCoDe: Boosting Model Performance Beyond Training with Extrapolation and Contrastive Decoding
- LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
- Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation
- Crowd-SFT: Crowdsourcing for LLM Alignment
- Do Large Language Models Know Folktales? A Case Study of Yokai in Japanese Folktales
- RewardAnything: Generalizable Principle-Following Reward Models
- MiMo-VL Technical Report
- Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
- Watermarking Degrades Alignment in Language Models: Analysis and Mitigation
- Multi-objective Aligned Bidword Generation Model for E-commerce Search Advertising
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
- SAGE:Specification-Aware Grammar Extraction for Automated Test Case Generation with LLMs
- Autonomous chemical research with large language models
- Robustness of Prompting: Enhancing Robustness of Large Language Models Against Prompting Attacks
- Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
- Protein Inverse Folding From Structure Feedback
- Geospatial Mechanistic Interpretability of Large Language Models
- Mutation-Guided Unit Test Generation with a Large Language Model
- World Modelling Improves Language Model Agents
- Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
- KARE-RAG: Knowledge-Aware Refinement and Enhancement for RAG
- Minos: A Multimodal Evaluation Model for Bidirectional Generation Between Image and Text
- BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflage
- XToM: Exploring the Multilingual Theory of Mind for Large Language Models
- ReSpace: Text-Driven 3D Indoor Scene Synthesis and Editing with Preference Alignment
- Should LLM Safety Be More Than Refusing Harmful Instructions?
- From Anger to Joy: How Nationality Personas Shape Emotion Attribution in Large Language Models
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
- IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages
- BNPO: Beta Normalization Policy Optimization
- Understanding the Impact of Sampling Quality in Direct Preference Optimization
- HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
- EvaLearn: Quantifying the Learning Capability and Efficiency of LLMs via Sequential Problem Solving
- Learning Together to Perform Better: Teaching Small-Scale LLMs to Collaborate via Preferential Rationale Tuning
- Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation Detection
- Truly Assessing Fluid Intelligence of Large Language Models through Dynamic Reasoning Evaluation
- Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences
- VerificAgent: Domain-Specific Memory Verification for Scalable Oversight of Aligned Computer-Use Agents
- AUTOCIRCUIT-RL: Reinforcement Learning-Driven LLM for Automated Circuit Topology Generation
- FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models
- MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching
- Seeing the Arrow of Time in Large Multimodal Models
- Rethinking Dynamic Networks and Heterogeneous Computing with Automatic Parallelization
- Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
- One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL
- Adversarial Attacks on Robotic Vision Language Action Models
- Beyond Text Compression: Evaluating Tokenizers Across Scales
- Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
- DPO Learning with LLMs-Judge Signal for Computer Use Agents
- EgoVLM: Policy Optimization for Egocentric Video Understanding
- Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
- Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
- EssayBench: Evaluating Large Language Models in Multi-Genre Chinese Essay Writing
- MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
- LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback
- What Makes LLMs Effective Sequential Recommenders? A Study on Preference Intensity and Temporal Context
- Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
- IF-GUIDE: Influence Function-Guided Detoxification of LLMs
- Can Large Language Models Provide Useful Feedback on Research Papers? A Large-Scale Empirical Analysis
- Agentic Episodic Control
- AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning
- Rating Quality of Diverse Time Series Data by Meta-learning from LLM Judgment
- Detoxification of Large Language Models through Output-layer Fusion with a Calibration Model
- CoRE: Condition-based Reasoning for Identifying Outcome Variance in Complex Events
- KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
- Incentivizing LLMs to Self-Verify Their Answers
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
- Stochastically Dominant Peer Prediction
- Synthetic Data Augmentation using Pre-trained Diffusion Models for Long-tailed Food Image Classification
- Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents
- When LLMs Team Up: The Emergence of Collaborative Affective Computing
- Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
- Benford's Curse: Tracing Digit Bias to Numerical Hallucination in LLMs
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
- Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?
- A Descriptive and Normative Theory of Human Beliefs in RLHF
- ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
- Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
- Fodor and Pylyshyn's Legacy: Still No Human-like Systematic Compositionality in Neural Networks
- Beyond RLHF: A Unified Theoretical Framework of Alignment
- RewardBench 2: Advancing Reward Model Evaluation
- Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner
- HASHIRU: Hierarchical Agent System for Hybrid Intelligent Resource Utilization
- The Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process
- XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content
- ACCESS DENIED INC: The First Benchmark Environment for Sensitivity Awareness
- Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language Models
- CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-Tuning
- Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
- Toward Structured Knowledge Reasoning: Contrastive Retrieval-Augmented Generation on Experience
- SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
- Improving LLM Agents with Reinforcement Learning on Cryptographic CTF Challenges
- Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge
- Doubly Robust Alignment for Large Language Models
- Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language Models
- Uni-LoRA: One Vector is All You Need
- Preference-based learning for news headline recommendation
- Scaling Textual Gradients via Sampling-Based Momentum
- CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries
- MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
- AutoMixAlign: Adaptive Data Mixing for Multi-Task Preference Optimization in LLMs
- Existing Large Language Model Unlearning Evaluations Are Inconclusive
- SafeTuneBed: A Toolkit for Benchmarking LLM Safety Alignment in Fine-Tuning
- GuideX: Guided Synthetic Data Generation for Zero-Shot Information Extraction
- Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs
- Alignment Revisited: Are Large Language Models Consistent in Stated and Revealed Preferences?
- Central Path Proximal Policy Optimization
- Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively
- SHARE: An SLM-based Hierarchical Action CorREction Assistant for Text-to-SQL
- QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
- RLAE: Reinforcement Learning-Assisted Ensemble for LLMs
- Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences
- VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making
- Soft Best-of-n Sampling for Model Alignment
- Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
- Optimizing Question Semantic Space for Dynamic Retrieval-Augmented Multi-hop Question Answering
- On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
- Beyond Multiple Choice: Evaluating Steering Vectors for Summarization
- From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning
- On Early Detection of Hallucinations in Factual Question Answering
- Adversarial Preference Learning for Robust LLM Alignment
- A Reward-driven Automated Webshell Malicious-code Generator for Red-teaming
- Benchmarking Foundation Models for Zero-Shot Biometric Tasks
- Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap
- RAST: Reasoning Activation in LLMs via Small-model Transfer
- AMSbench: A Comprehensive Benchmark for Evaluating MLLM Capabilities in AMS Circuits
- Reasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning Models
- Aligning Protein Conformation Ensemble Generation with Physical Feedback
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective
- Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
- Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning
- MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning
- Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise
- Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
- Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning
- Intuitionistic Fuzzy Sets for Large Language Model Data Annotation: A Novel Approach to Side-by-Side Preference Labeling
- ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
- MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
- CRScore++: Reinforcement Learning with Verifiable Tool and AI Feedback for Code Review
- Who Gets Credit or Blame? Attributing Accountability in Modern AI Systems
- Proxy Target: Bridging the Gap Between Discrete Spiking Neural Networks and Continuous Control
- Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation
- Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
- Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment
- Tag-Evol: Achieving Efficient Instruction Evolving via Tag Injection
- When Large Multimodal Models Confront Evolving Knowledge: Challenges and Explorations
- A Red Teaming Roadmap Towards System-Level Safety
- DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
- The End Of Universal Lifelong Identifiers: Identity Systems For The AI Era
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- BIRD: Behavior Induction via Representation-structure Distillation
- Thompson Sampling in Online RLHF with General Function Approximation
- Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
- LLMs are Better Than You Think: Label-Guided In-Context Learning for Named Entity Recognition
- Grounded Reinforcement Learning for Visual Reasoning
- SLOT: Structuring the Output of Large Language Models
- PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward Model
- Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
- Fortune: Formula-Driven Reinforcement Learning for Symbolic Table Reasoning in Language Models
- SafeScientist: Toward Risk-Aware Scientific Discoveries by LLM Agents
- Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information
- Learning Parametric Distributions from Samples and Preferences
- Probability-Consistent Preference Optimization for Enhanced LLM Reasoning
- Identity resolution of software metadata using Large Language Models
- Mis-prompt: Benchmarking Large Language Models for Proactive Error Handling
- Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
- Cross-Task Experiential Learning on LLM-based Multi-Agent Collaboration
- MARCO: Multi-Agent Code Optimization with Real-Time Knowledge Integration for High-Performance Computing
- PhotoArtAgent: Intelligent Photo Retouching with Language Model-Based Artist Agents
- Dataset Cartography for Large Language Model Alignment: Mapping and Diagnosing Preference Data
- LLM Agents for Bargaining with Utility-based Feedback
- Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios
- OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Languages and Modalities
- Accelerating RLHF Training with Reward Variance Increase
- Iterative Resolution of Prompt Ambiguities Using a Progressive Cutting-Search Approach
- Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models
- How Does Response Length Affect Long-Form Factuality
- Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models
- Document-Level Text Generation with Minimum Bayes Risk Decoding using Optimal Transport
- Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO
- Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
- Neither Stochastic Parroting nor AGI: LLMs Solve Tasks through Context-Directed Extrapolation from Training Data Priors
- Bigger, Regularized, Categorical: High-Capacity Value Functions are Efficient Multi-Task Learners
- MenTeR: A fully-automated Multi-agenT workflow for end-to-end RF/Analog Circuits Netlist Design
- SC-LoRA: Balancing Efficient Fine-tuning and Knowledge Preservation via Subspace-Constrained LoRA
- Continuous Chain of Thought Enables Parallel Exploration and Reasoning
- LlamaRL: A Distributed Asynchronous Reinforcement Learning Framework for Efficient Large-scale LLM Training
- LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generations
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
- On-Policy RL with Optimal Reward Baseline
- DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
- Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
- Understanding Refusal in Language Models with Sparse Autoencoders
- MAP: Revisiting Weight Decomposition for Low-Rank Adaptation
- Are Reasoning Models More Prone to Hallucination?
- Operationalizing CaMeL: Strengthening LLM Defenses for Enterprise Deployment
- Preference Learning with Response Time: Robust Losses and Guarantees
- MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
- Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation
- Measuring Sycophancy of Language Models in Multi-turn Dialogues
- Pangu Embedded: An Efficient Dual-system LLM Reasoner with Metacognition
- GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance
- Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments
- Align-DA: Align Score-based Atmospheric Data Assimilation with Multiple Preferences
- Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy
- SDPO: Importance-Sampled Direct Preference Optimization for Stable Diffusion Training
- Reinforced Reasoning for Embodied Planning
- MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models
- Decomposing Elements of Problem Solving: What "Math" Does RL Teach?
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
- Zero-Shot 3D Visual Grounding from Vision-Language Models
- ImageReFL: Balancing Quality and Diversity in Human-Aligned Diffusion Models
- Seven Security Challenges in Cross-domain Multi-agent LLM Systems
- A Provable Approach for End-to-End Safe Reinforcement Learning
- LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agents
- Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development
- BiasFilter: An Inference-Time Debiasing Framework for Large Language Models
- Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
- Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
- From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications
- Xinyu AI Search: Enhanced Relevance and Comprehensive Results with Rich Answer Presentations
- Fostering Video Reasoning via Next-Event Prediction
- Beyond path selection: Better LLMs for Scientific Information Extraction with MimicSFT and Relevance and Rule-induced(R2)GRPO
- RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
- Text2Grad: Reinforcement Learning from Natural Language Feedback
- ArgInstruct: Specialized Instruction Fine-Tuning for Computational Argumentation
- Communication-Efficient Desire Alignment for Embodied Agent-Human Adaptation
- EvolveSearch: An Iterative Self-Evolving Search Agent
- Rethinking the Unsolvable: When In-Context Search Meets Test-Time Scaling
- D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples
- 360-LLaMA-Factory: Plug & Play Sequence Parallelism for Long Post-Training
- The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason
- Reverse Preference Optimization for Complex Instruction Following
- Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation
- Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training
- LASER: Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Towards Safety Reasoning in LLMs: AI-agentic Deliberation for Policy-embedded CoT Data Creation
- Relevance-driven Input Dropout: an Explanation-guided Regularization Technique
- Leveraging Large Language Models for Bengali Math Word Problem Solving with Chain of Thought Reasoning
- Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
- Who Reasons in the Large Language Models?
- Concealment of Intent: A Game-Theoretic Analysis
- SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
- RRO: LLM Agent Optimization Through Rising Reward Trajectories
- MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
- Uni-Instruct: One-step Diffusion Model through Unified Diffusion Divergence Instruction
- Reinforcing General Reasoning without Verifiers
- SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation
- STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
- Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models
- AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
- LPOI: Listwise Preference Optimization for Vision Language Models
- Knowing You Don't Know: Learning When to Continue Search in Multi-round RAG through Self-Practicing
- Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching
- Personalized Query Auto-Completion for Long and Short-Term Interests with Adaptive Detoxification Generation
- Towards Better Instruction Following Retrieval Models
- MIRROR: Multi-agent Intra- and Inter-Reflection for Optimized Reasoning in Tool Learning
- Policy Optimized Text-to-Image Pipeline Design
- Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
- Large Language Models Miss the Multi-Agent Mark
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment
- A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs
- SquareχPO: Differentially Private and Robust χ2-Preference Optimization in Offline Direct Alignment
- Contrastive Learning on LLM Back Generation Treebank for Cross-domain Constituency Parsing
- DreamBoothDPO: Improving Personalized Generation using Direct Preference Optimization
- IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios
- RM-R1: Reward Modeling as Reasoning
- Aligning LLMs by Predicting Preferences from User Writing Samples
- Explaining Large Language Models with gSMILE
- Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
- SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
- Pretrained LLMs Learn Multiple Types of Uncertainty
- MultiPhishGuard: An LLM-based Multi-Agent System for Phishing Email Detection
- Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process Supervision
- Incentivizing Inclusive Contributions in Model Sharing Markets
- Temporal Sampling for Forgotten Reasoning in LLMs
- Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
- What Can RL Bring to VLA Generalization? An Empirical Study
- T2Agent A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree Search
- Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models
- Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
- Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models
- Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
- Prompting is not Enough: Exploring Knowledge Integration and Controllable Generation
- Token-Importance Guided Direct Preference Optimization
- Risk-aware Direct Preference Optimization under Nested Risk Measure
- JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models
- Energy-based Preference Optimization for Test-time Adaptation
- Preference Optimization by Estimating the Ratio of the Data Distribution
- Learning to Reason without External Rewards
- Learning to Select In-Context Demonstration Preferred by Large Language Model
- StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
- Self-Reflective Planning with Knowledge Graphs: Enhancing LLM Reasoning Reliability for Question Answering
- Prot2Token: A Unified Framework for Protein Modeling via Next-Token Prediction
- The Limits of Preference Data for Post-Training
- Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries
- Learning a Pessimistic Reward Model in RLHF
- Inference-time Alignment in Continuous Space
- Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
- CPA-RAG:Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models
- Interleaved Reasoning for Large Language Models via Reinforcement Learning
- Alignment of large language models with constrained learning
- S2LPP: Small-to-Large Prompt Prediction across LLMs
- RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback
- CaseEdit: Enhancing Localized Commonsense Reasoning via Null-Space Constrained Knowledge Editing in Small Parameter Language Models
- Does Rationale Quality Matter? Enhancing Mental Disorder Detection via Selective Reasoning Distillation
- Toward Scientific Reasoning in LLMs: Training from Expert Discussions via Reinforcement Learning
- Conversation Kernels: A Flexible Mechanism to Learn Relevant Context for Online Conversation Understanding
- FunReason: Enhancing Large Language Models' Function Calling via Self-Refinement Multiscale Loss and Automated Data Refinement
- On the Same Page: Dimensions of Perceived Shared Understanding in Human-AI Interaction
- FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering
- Opt-GPTQ: An Optimized GPTQ Combining Sparse Attention and Quantization Techniques
- SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
- FairPO: Robust Preference Optimization for Fair Multi-Label Learning
- Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
- Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play
- Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
- Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
- SCAR: Shapley Credit Assignment for More Efficient RLHF
- SIPDO: Closed-Loop Prompt Optimization via Synthetic Data Feedback
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
- Safety Through Reasoning: An Empirical Study of Reasoning Guardrail Models
- SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
- Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
- ARM: Adaptive Reasoning Model
- Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
- Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers
- Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
- EGA-V1: Unifying Online Advertising with End-to-End Learning
- A Survey on Progress in LLM Alignment from the Perspective of Reward Design
- Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents
- What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMs
- Improving Value Estimation Critically Enhances Vanilla Policy Gradient
- ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
- LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
- Incentivizing High-Quality Human Annotations with Golden Questions
- An Embarrassingly Simple Defense Against LLM Abliteration Attacks
- Online Knowledge Distillation with Reward Guidance
- A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement Learning
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems
- The Price of Format: Diversity Collapse in LLMs
- Do Large Language Models (Really) Need Statistical Foundations?
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts
- VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
- Behavior Injection: Preparing Language Models for Reinforcement Learning
- Mitigating Deceptive Alignment via Self-Monitoring
- Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
- AI-Driven Climate Policy Scenario Generation for Sub-Saharan Africa
- Benchmarking and Rethinking Knowledge Editing for Large Language Models
- Language Model Distillation: A Temporal Difference Imitation Learning Perspective
- Robustness in Large Language Models: A Survey of Mitigation Strategies and Evaluation Metrics
- Skip-Thinking: Chunk-wise Chain-of-Thought Distillation Enable Smaller Language Models to Reason Better and Faster
- Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
- Unraveling Misinformation Propagation in LLM Reasoning
- Diffusion Blend: Inference-Time Multi-Preference Alignment for Diffusion Models
- Post-Training on Office Work Improves Software Engineering: A Behavioral Account of Cross-Domain Transfer
- Synthesizing and Adapting Error Correction Data for Mobile Large Language Model Applications
- TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
- The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
- VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
- From Generation to Detection: A Multimodal Multi-Task Dataset for Benchmarking Health Misinformation
- Safety Alignment via Constrained Knowledge Unlearning
- Hybrid Latent Reasoning via Reinforcement Learning
- Rethinking Direct Preference Optimization in Diffusion Models
- Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs
- PD3F: A Pluggable and Dynamic DoS-Defense Framework Against Resource Consumption Attacks Targeting Large Language Models
- Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models
- Generative RLHF-V: Learning Principles from Multi-modal Human Preference
- MOSLIM:Align with diverse preferences in prompts through reward classification
- PromptWise: Online Learning for Cost-Aware Prompt Assignment in Generative Models
- OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
- Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
- Test-Time Adaptation with Binary Feedback
- Advertising in AI systems: Society must be vigilant
- Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
- Reward Model Overoptimisation in Iterated RLHF
- Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
- Stable Reinforcement Learning for Efficient Reasoning
- Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals
- MathEDU: Towards Adaptive Feedback for Student Mathematical Problem-Solving
- Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL
- Scalable Valuation of Human Feedback through Provably Robust Model Alignment
- ChatGPT: Jack of all trades, master of none
- DialogXpert: Driving Intelligent and Emotion-Aware Conversations through Online Value-Based Reinforcement Learning with LLM Priors
- Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning
- Self-Improving Large Language Models via Progressive Experience Evolution
- Understanding How Value Neurons Shape the Generation of Specified Values in LLMs
- Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models
- ViP2-CLIP: Visual-Perception Prompting with Unified Alignment for Zero-Shot Anomaly Detection
- Tuning Language Models for Robust Prediction of Diverse User Behaviors
- Emergent Standing Wave Dynamics and Attractor Basins in Transformer Latent Spaces via Prompt Driven ConstraintsAnathema to Corporate Control by Ingrid Johnson
- Start Classifying: Categorical Critics for LLM Reinforcement Learning
- One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
- JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
- Universal Biological Sequence Reranking for Improved De Novo Peptide Sequencing
- Large Language Models Do Multi-Label Classification Differently
- SLearnLLM: A Self-Learning Framework for Efficient Domain-Specific Adaptation of Large Language Models
- Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
- Curriculum Guided Reinforcement Learning for Efficient Multi Hop Retrieval Augmented Generation
- AI-Augmented LLMs Achieve Therapist-Level Responses in Motivational Interviewing
- LA-RCS: LLM-Agent-Based Robot Control System
- Alignment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey
- On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
- Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
- Trinity-RFT: A General-Purpose and Unified Framework for Reinforcement Fine-Tuning of Large Language Models
- Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary
- KL-regularization Itself is Differentially Private in Bandits and RLHF
- Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
- Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling
- Arctic-Text2SQL-R1: Simple Rewards, Strong Reasoning in Text-to-SQL
- Refusal Direction is Universal Across Safety-Aligned Languages
- Learning to Choose or Choosing to Learn: Best-of-N vs. Supervised Fine-Tuning for Bit String Generation
- ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
- ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialects
- Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts
- Effective Reinforcement Learning for Reasoning in Language Models
- Shape it Up! Restoring LLM Safety during Finetuning
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
- MPO: Multilingual Safety Alignment via Reward Gap Optimization
- BitHydra: Towards Bit-flip Inference Cost Attack against Large Language Models
- Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
- CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
- SC4ANM: Identifying Optimal Section Combinations for Automated Novelty Prediction in Academic Papers
- Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
- Offline Guarded Safe Reinforcement Learning for Medical Treatment Optimization Strategies
- LightRouter: Towards Efficient LLM Collaboration with Minimal Overhead
- LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
- Dynamic Sampling that Adapts: Iterative DPO for Self-Aware Mathematical Reasoning
- RLKD: Distilling LLMs' Reasoning via Reinforcement Learning
- Your Pre-trained LLM is Secretly an Unsupervised Confidence Calibrator
- MARché: Fast Masked Autoregressive Image Generation with Cache-Aware Attention
- Sudoku-Bench: Evaluating creative reasoning with Sudoku variants
- Walk&Retrieve: Simple Yet Effective Zero-shot Retrieval-Augmented Generation via Knowledge Graph Walks
- SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development
- Data-Driven Breakthroughs and Future Directions in AI Infrastructure: A Comprehensive Review
- Action is All You Need: Dual-Flow Generative Ranking Network for Recommendation
- R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
- CASTILLO: Characterizing Response Length Distributions of Large Language Models
- SATURN: SAT-based Reinforcement Learning to Unleash LLMs Reasoning
- Interactive Post-Training for Vision-Language-Action Models
- ReCopilot: Reverse Engineering Copilot in Binary Analysis
- LLM Fingerprinting via Semantically Conditioned Watermarks
- MPL: Multiple Programming Languages with Large Language Models for Information Extraction
- Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models
- SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
- Understanding Fact Recall in Language Models: Why Two-Stage Training Encourages Memorization but Mixed Training Teaches Knowledge
- Serious Games: Human-AI Interaction, Evolution, and Coevolution
- Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs
- DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimization
- Recursive Offloading for LLM Serving in Multi-tier Networks
- From Generic Empathy to Personalized Emotional Support: A Self-Evolution Framework for User Preference Alignment
- Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning
- Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
- BiasLab: Toward Explainable Political Bias Detection with Dual-Axis Annotations and Rationale Indicators
- Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
- Explaining Puzzle Solutions in Natural Language: An Exploratory Study on 6x6 Sudoku
- CoT Information: Improved Sample Complexity under Chain-of-Thought Supervision
- Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data?
- Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
- Direct Preference Optimization for Adaptive Concept-based Explanations
- Joint Flashback Adaptation for Forgetting-Resistant Instruction Tuning
- Emotional Supporters often Use Multiple Strategies in a Single Turn
- Cross-Domain Hybrid OPD for Generalizable Search Agents
- Kernel PCA for Out-of-Distribution Detection: Non-Linear Kernel Selections and Approximations
- Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
- Flashback: Memory-Driven Zero-shot, Real-time Video Anomaly Detection
- AvatarShield: Visual Reinforcement Learning for Human-Centric Synthetic Video Detection
- DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
- Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
- Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment
- The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
- A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
- Set-LLM: A Permutation-Invariant LLM
- Perception-Driven Bias Detection in Machine Learning via Crowdsourced Visual Judgment
- From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning
- CoLA: Collaborative Low-Rank Adaptation
- Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains
- Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
- Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
- Position: Agentic Systems Constitute a Key Component of Next-Generation Intelligent Image Processing
- Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
- LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models
- VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
- Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition
- Reward Is Enough: LLMs Are In-Context Reinforcement Learners
- Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
- KaFT: Knowledge-aware Fine-tuning for Boosting LLMs' Domain-specific Question-Answering Performance
- Gated Integration of Low-Rank Adaptation for Continual Learning of Large Language Models
- An Empirical Study on Reinforcement Learning for Reasoning-Search Interleaved LLM Agents
- Learning from Algorithm Feedback: One-Shot SAT Solver Guidance with GNNs
- Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
- Few-Shot Test-Time Optimization Without Retraining for Semiconductor Recipe Generation and Beyond
- Emerging Properties in Unified Multimodal Pretraining
- NExT-Search: Rebuilding User Feedback Ecosystem for Generative AI Search
- Advancing Multi-Agent RAG Systems with Minimalist Reinforcement Learning
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
- Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
- sudoLLM: On Multi-role Alignment of Language Models
- Context Reasoner: Incentivizing Reasoning Capability for Contextualized Privacy and Safety Compliance via Reinforcement Learning
- QA-prompting: Improving Summarization with Large Language Models using Question-Answering
- YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
- AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Momentum
- Safety Degradation in AI Agents
- RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
- Think-J: Learning to Think for Generative LLM-as-a-Judge
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- MonitorVLM-v2: A Deployed Vision-Language Framework for Real-Time Safety Violation Detection
- Creative Preference Optimization
- Fragments to Facts: Partial-Information Fragment Inference from LLMs
- Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
- MCIP: Protecting MCP Safety via Model Contextual Integrity Protocol
- RLVR-World: Training World Models with Reinforcement Learning
- General-Reasoner: Advancing LLM Reasoning Across All Domains
- SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks
- Visual Agentic Reinforcement Fine-Tuning
- Debating for Better Reasoning: An Unsupervised Multimodal Approach
- Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art
- Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
- Self-Evolving Curriculum for LLM Reasoning
- AUTOLAW: Enhancing Legal Compliance in Large Language Models via Case Law Generation and Jury-Inspired Deliberation
- Reward Reasoning Model
- Reliable Decision Support with LLMs: A Framework for Evaluating Consistency in Binary Text Classification Applications
- Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
- PandaGuard: Systematic Evaluation of LLM Safety against Jailbreaking Attacks
- Domain Gating Ensemble Networks for AI-Generated Text Detection
- Reinforcement Learning from User Feedback
- FisherSFT: Data-Efficient Supervised Fine-Tuning of Language Models Using Information Gain
- Towards eliciting latent knowledge from LLMs with mechanistic interpretability
- Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency
- SQLForge: Synthesizing Reliable and Diverse Data to Enhance Text-to-SQL Reasoning in LLMs
- Safety Alignment Can Be Not Superficial With Explicit Safety Signals
- Metric Distortion for Tournament Voting and Beyond
- CIE: Controlling Language Model Text Generations Using Continuous Signals
- Fine-tuning Quantized Neural Networks with Zeroth-order Optimization
- What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
- J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization
- Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
- Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning
- Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs
- DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management
- Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks
- Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
- PsyMem: Fine-grained psychological alignment and Explicit Memory Control for Advanced Role-Playing LLMs
- Combining the Best of Both Worlds: A Method for Hybrid NMT and LLM Translation
- Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
- PromptPrism: A Linguistically-Inspired Taxonomy for Prompts
- Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
- A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
- Walking the Tightrope: Disentangling Beneficial and Detrimental Drifts in Non-Stationary Custom-Tuning
- Reasoning BO: Enhancing Bayesian Optimization with Long-Context Reasoning Power of LLMs
- Learnware of Language Models: Specialized Small Language Models Can Do Big
- MR. Judge: Multimodal Reasoner as a Judge
- Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
- Is Active Persona Inference Necessary for Aligning Small Models to Personal Preferences?
- Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks
- Do people rely on ChatGPT more than their peers to detect deepfake news?
- When Retrieval Helps and Distracts: Evaluating Evidence-Generating LLMs for Biomedical Claim Verification
- Reusing Rollouts under Policy Lag: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning
- Fractured Chain-of-Thought Reasoning
- CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval
- SayCoNav: Utilizing Large Language Models for Adaptive Collaboration in Decentralized Multi-Robot Navigation
- R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
- Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
- Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
- CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process
- Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
- Adaptive Tokenization: On the Hop-Overpriority Problem in Tokenized Graph Learning Models
- Krikri: Advancing Open Large Language Models for Greek
- Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks
- Bullying the Machine: How Personas Increase LLM Vulnerability
- ChemPile: A 250GB Diverse and Curated Dataset for Chemical Foundation Models
- Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration
- Towards DS-NER: Unveiling and Addressing Latent Noise in Distant Annotations
- BadNAVer: Exploring Jailbreak Attacks On Vision-and-Language Navigation
- Knowledge graph–based thought: a knowledge graph–enhanced LLM framework for pan-cancer question answering
- The Role of ChatGPT and AI Chatbots in Optimizing Antibiotic Therapy: A Comprehensive Narrative Review
- Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments
- Bridging Generative and Discriminative Learning: Few-Shot Relation Extraction via Two-Stage Knowledge-Guided Pre-training
- Data Whisperer: Efficient Data Selection for Task-Specific LLM Fine-Tuning via Few-Shot In-Context Learning
- BrainOmni: A Brain Foundation Model for Unified EEG and MEG Signals
- Enriching Patent Claim Generation with European Patent Dataset
- DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization
- NeuroGen: Neural Network Parameter Generation via Large Language Models
- EVALOOOP: A Self-Consistency-Centered Framework for Assessing Large Language Model Robustness in Programming
- Self-Destructive Language Model
- LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
- SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment
- Not All Documents Are What You Need for Extracting Instruction Tuning Data
- SPIRIT: Patching Speech Language Models against Jailbreak Attacks
- SchoenbAt: Rethinking Attention with Polynomial basis
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
- Beyond Single-Point Judgment: Distribution Alignment for LLM-as-a-Judge
- AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
- Pairwise Calibrated Rewards for Pluralistic Alignment
- Spotlight Your Instructions: Instruction-following with Dynamic Attention Steering
- Security practices in AI development
- Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier
- SafeVid: Toward Safety Aligned Video Large Multimodal Models
- Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform
- VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
- Multilingual Collaborative Defense for Large Language Models
- When Do Surrogate Updates Improve Decisions? A Local Theory of Trajectory-Wise Transfer
- AdaCoT: Pareto-Optimal Adaptive Chain-of-Thought Triggering via Reinforcement Learning
- Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
- Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning
- Solver-Informed RL: Grounding Large Language Models for Authentic Optimization Modeling
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
- Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Reward Design
- Enhancing Complex Instruction Following for Large Language Models with Mixture-of-Contexts Fine-tuning
- Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
- AdaBoN: Adaptive Best-of-N Alignment
- Mutual-Taught for Co-adapting Policy and Reward Models
- J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
- Stepwise Guided Policy Optimization: Coloring your Incorrect Reasoning in GRPO
- Visual Planning: Let's Think Only with Images
- Time-R1: Towards Comprehensive Temporal Reasoning in LLMs
- Is PRM Necessary? Problem-Solving RL Implicitly Induces PRM Capability in LLMs
- Hierarchical Residual Policy Optimization for Generative Recommendations
- Lexicalization Is All You Need: Examining the Impact of Lexical Knowledge in a Compositional QALD System
- DiffuseAgent-MI: Distributionally-Grounded,Tool-Integrated Self-Evolving Agents for Faithful Visual Reasoning
- Where did the ambiguity go? Examining how multimodal models interpret polysemous words
- AiFlow: Token-Native Reactive Orchestration with Bounded Backpressure for Streaming LLM Applications
- InstructPLM: Aligning Protein Language Models to Follow Protein Structure Instructions
- Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations
- Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
- Attention-Based Reward Shaping for Sparse and Delayed Rewards
- MergeBench: A Benchmark for Merging Domain-Specialized LLMs
- ShiQ: Bringing back Bellman to LLMs
- ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training
- Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
- Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
- Illusions of Intimacy: How Emotional Dynamics Shape Human-AI Relationships
- LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
- REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning
- GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
- Finetune-RAG: Fine-Tuning Language Models to Resist Hallucination in Retrieval-Augmented Generation
- Learning from Less: Guiding Deep Reinforcement Learning with Differentiable Symbolic Planning
- Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
- A Systematic Analysis of Base Model Choice for Reward Modeling
- Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback
- GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
- Group-in-Group Policy Optimization for LLM Agent Training
- SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance
- Feasibility with Language Models for Open-World Compositional Zero-Shot Learning
- BLEUBERI: BLEU is a surprisingly effective reward for instruction following
- Diffusion-NPO: Negative Preference Optimization for Better Preference Aligned Generation of Diffusion Models
- When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
- Ranked Voting based Self-Consistency of Large Language Models
- SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization
- Cochain: Balancing Insufficient and Excessive Collaboration in LLM Agent Workflows
- Review-Instruct: A Review-Driven Multi-Turn Conversations Generation Method for Large Language Models
- Who You Are Matters: Bridging Topics and Social Roles via LLM-Enhanced Logical Recommendation
- Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented Reasoning
- A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
- Generative Muscle Stimulation: Physical Assistance by Constraining Multimodal-AI with Biomechanical Knowledge
- WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
- Auditable Release Control for Pedagogical Leakage in LLM Tutors
- Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
- Pre-Act: Multi-Step Planning and Reasoning Improves Acting in LLM Agents
- PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
- T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
- Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language Models
- Demystifying AI Agents: The Final Generation of Intelligence
- RSRA: Training-Free Probing of Representation Sensitivity for Efficient LoRA Rank Allocation
- Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
- Interpretable Risk Mitigation in LLM Agent Systems
- WorldPM: Scaling Human Preference Modeling
- TAB-PO: Preference Optimization with a Token-Level Adaptive Barrier for Token-Critical Structured Generation
- Distributionally Robust Listwise Preference Optimization
- Benchmarking LLMs on File System Design and Implementation
- Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning
- Reasoning Capabilities of Large Language Models on Dynamic Tasks
- J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
- Inference-Time Policy Alignment for Fair Reinforcement Learning
- DRA-GRPO: Exploring Diversity-Aware Reward Adjustment for R1-Zero-Like Training of Large Language Models
- Atomic Consistency Preference Optimization for Long-Form Question Answering
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning
- InvDesFlow-AL: active learning-based workflow for inverse design of functional materials
- Ethics and Persuasion in Reinforcement Learning from Human Feedback: A Procedural Rhetorical Approach
- LLM4CD: Leveraging Large Language Models for Open-World Knowledge Augmented Cognitive Diagnosis
- System Prompt Optimization with Meta-Learning
- Language Agents Mirror Human Causal Reasoning Biases. How Can We Help Them Think Like Scientists?
- Verifier-Induced Support Reshaping in On-Policy Optimization
- Memorization-Compression Cycles Improve Generalization
- Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
- Large Language Models Meet Stance Detection: A Survey of Tasks, Methods, Applications, Challenges and Future Directions
- TUMS: Enhancing Tool-use Abilities of LLMs with Multi-structure Handlers
- GenPage: Towards End-to-End Generative Homepage Construction at Netflix
- Evaluating LLM Metrics Through Real-World Capabilities
- Freeform Preference Learning for Robotic Manipulation
- Improved Algorithms for Differentially Private Language Model Alignment
- Memory Reward Inflation in Self-Improving LLM Agents
- Detecting Prefix Bias in LLM-based Reward Models
- Large Language Models for Computer-Aided Design: A Survey
- Assessing and Mitigating Medical Knowledge Drift and Conflicts in Large Language Models
- S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models
- A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models
- BinMetric: A Comprehensive Binary Analysis Benchmark for Large Language Models
- On the Robustness of Reward Models for Language Model Alignment
- DanceGRPO: Unleashing GRPO on Visual Generation
- Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
- Concept-Level Explainability for Auditing & Steering LLM Responses
- Kalman Filter Enhanced GRPO for Reinforcement Learning-Based Language Model Reasoning
- Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning
- You Only Look One Step: Accelerating Backpropagation in Diffusion Sampling with Gradient Shortcuts
- DynamicRAG: Leveraging Outputs of Large Language Model as Feedback for Dynamic Reranking in Retrieval-Augmented Generation
- TrumorGPT: Graph-Based Retrieval-Augmented Large Language Model for Fact-Checking
- A Survey on Foundation Models for Personalized Federated Intelligence
- PLHF: Prompt Optimization with Few-Shot Human Feedback
- Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks Safety
- Learning Guarantee of Reward Modeling Using Deep Neural Networks
- REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback
- Towards Embodiment Scaling Laws in Robot Locomotion
- A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook
- (How) Learning Rates Regulate Catastrophic Overtraining
- GeoMind: An Agentic Workflow for Lithology Classification with Reasoned Tool Invocation
- Towards Developmentally Plausible Rewards: Communicative Success as a Learning Signal for Interactive Language Models
- Neural Catalog: Scaling Species Recognition with Catalog of Life-Augmented Generation
- MAPLE: Metadata Augmented Private Language Evolution
- Scaling Laws for Speculative Decoding
- Multi-agent Embodied AI: Advances and Future Directions
- Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes
- Tail-Aware Information-Theoretic Bounds for LLM Alignment under Heavy-Tailed Rewards
- ComPO: Preference Alignment via Comparison Oracles
- Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
- SVRepair: Structured Visual Reasoning for Automated Program Repair
- Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation
- MedTextWeaver: Procedural Knowledge Evolution in Agentic Medical Text Editing
- Any House Any Task: Scalable Long-Horizon Planning for Abstract Human Tasks
- DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference
- CORE: Collaborative Reasoning via Cross Teaching
- When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
- A Survey on Privacy Risks and Protection in Large Language Models
- Demystifying optimized prompts in language models
- Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data
- Adaptive Social Learning via Mode Policy Optimization for Language Agents
- R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
- Exploring the Potential of Offline RL for Reasoning in LLMs: A Preliminary Study
- Semantic Probabilistic Control of Language Models
- AI-Based Speaking Assistant: Supporting Non-Native Speakers' Speaking in Real-Time Multilingual Communication
- LookAlike: Consistent Distractor Generation in Math MCQs
- Reinforcement Learning for Speculative Trading under Exploratory Framework
- Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
- Contextures: Representations from Contexts
- Suffix-Constrained Greedy Search Algorithms for Causal Language Models
- Multi-agents based User Values Mining for Recommendation
- Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
- Transferable Adversarial Attacks on Black-Box Vision-Language Models
- Aligning Large Language Models with Healthcare Stakeholders: A Pathway to Trustworthy AI Integration
- Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models
- Steering Large Language Models with Register Analysis for Arbitrary Style Transfer
- DeepCritic: Deliberate Critique with Large Language Models
- FineScope : Precision Pruning for Domain-Specialized Large Language Models Using SAE-Guided Self-Data Cultivation
- Sustainability Is Not Linear: Quantifying Performance, Energy, and Privacy Trade-offs in On-Device Intelligence
- Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
- R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- DUPLEX: Agentic Dual-System Planning via LLM-Driven Information Extraction
- Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact
- RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
- Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
- Alignment midtraining for animals
- Safety in Batches? Understanding and Mitigating Safety Failures in Batch Prompting
- Quo Vadis, World Modeling?
- Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment
- Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks
- Teaching an Agent to Sketch One Part at a Time
- Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models
- Are We Ready for AI-Driven Discovery? AI Verification Before the Next Fundamental Physics Breakthrough
- Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
- Governable Individuals: An Identity Layer for Embodied Agents That Keep Learning
- When Context Returns: Toward Robust Internalization in On-Policy Distillation
- Attractor States Emerge in Multi-Turn LLM Conversations
- INSID3: Training-Free In-Context Segmentation with DINOv3
- Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth Retrieval
- 100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models
- Retrieval-Feedback-Driven Distillation and Preference Alignment for Efficient LLM-based Query Expansion
- QTrack: Query-Driven Reasoning for Multi-modal MOT
- Colluding LoRA: A Compositional Vulnerability in LLM Safety Alignment
- Generative Engine Optimization: A VLM and Agent Framework for Pinterest Acquisition Growth
- Enabling Cloud-Level Accuracy in Edge AI through IoT Data Preprocessing
- Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding
- Large Language Models Do Not Always Need Readable Language
- Cordon: Semantic Transactions for Tool-Using LLM Agents
- The Value Axis: Language Models Encode Whether They're on the Right Track
- Assessing Automated Prompt Injection Attacks in Agentic Environments
- GRPO Does Not Close the Multi-Agent Coordination Gap
- Multi-Stream Perturbation Attack: Breaking Safety Alignment of Thinking LLMs Through Concurrent Task Interference
- AlphaLogics: A Market Logic-Driven Multi-Agent System for Scalable and Interpretable Alpha Factor Generation
- The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
- Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
- Toward Epistemic Stability: Engineering Consistent Procedures for Industrial LLM Hallucination Reduction
- Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers
- What Do People Actually Want From AI? Mapping Preference Plurality
- Re-Centering Humans in LLM Personalization
- Reward-Conditioned Reinforcement Learning
- Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use
- Sleeper Cell: Injecting Latent Malice Temporal Backdoors into Tool-Using LLMs
- Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
- Jailbreaking Embodied LLMs via Action-level Manipulation
- Demystifying Group Relative Policy Optimization: Its Policy Gradient is a U-Statistic
- RAVEL: Reasoning Agents for Validating and Evaluating LLM Text Synthesis
- Foundation World Models for Agents that Learn, Verify, and Adapt Reliably Beyond Static Environments
- Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
- Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoning
- Requesting Expert Reasoning: Augmenting LLM Agents with Learned Collaborative Intervention
- BrepCoder: A Unified Multimodal Large Language Model for Multi-task B-rep Reasoning
- StepPRM-RTL: Stepwise Process-Reward Guided LLM Fine-Tuning for Enhanced RTL Synthesis
- Investigating Zero-Shot Diagnostic Pathology in Vision-Language Models with Efficient Prompt Design
- Power and Limitations of Aggregation in Compound AI Systems
- Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
- Grid-Mind: An LLM-Orchestrated Multi-Fidelity Agent for Automated Connection Impact Assessment
- Spilled Energy in Large Language Models
- Capabilities Ain't All You Need: Measuring Propensities in AI
- Creating a digital poet
- RIDER: 3D RNA Inverse Design with Reinforcement Learning-Guided Diffusion
- Beyond Refusal: A Same-Lineage Study of Aligned and Abliterated LLMs for Vulnerability Analysis
- Positive Alignment: Artificial Intelligence for Human Flourishing
- When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape
- Springdrift: An Auditable Persistent Runtime for LLM Agents with Case-Based Memory, Normative Safety, and Ambient Self-Perception
- The Last Fingerprint: How Markdown Training Shapes LLM Prose
- Progressive Training for Explainable Citation-Grounded Dialogue: Reducing Hallucination to Zero in English-Hindi LLMs
- Long-Tail Knowledge in Large Language Models: Taxonomy, Mechanisms, Interventions and Implications
- Intent Laundering: AI Safety Datasets Are Not What They Seem
- Engineering Reasoning and Instruction (ERI) Benchmark: A Large Taxonomy-driven Dataset for Foundation Models and Agents
- Chatbots Output Meaningful (but Problematic) Language
- SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
- Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search
- A New Framework for Cybersecurity Refusals in AI Agents
- Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling
- Configurable Reward Model for Balanced Safety Alignment
- Hierarchical Prompt-Domain Control and Learning for Resource-Constrained Agentic Language Models
- In LLM Reasoning, there is Irrationality on top of Value Misalignment
- Causal methods for LLM development and evaluation
- Insuring Every Action: An Authority Frontier Framework for Runtime Actuarial Control of Autonomous AI Agents
- Learning to Route Languages for Multilingual Policy Optimization
- Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
- JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment
- Look Inward to Explore Outward: Learning Temperature Policy from LLM Internal States via Hierarchical RL
- Using Large Language Models in Physics Education
- Trust The Typical
- A New Strategy for Artificial Intelligence: Training Foundation Models Directly on Human Brain Data
- XBreaking: Understanding how LLMs security alignment can be broken
- Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs' Instruction Following Capability
- Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications
- Confidence in Large Language Model Evaluation: A Bayesian Approach to Limited-Sample Challenges
- ShorterBetter: Guiding Reasoning Models to Find Optimal Inference Length for Efficient Reasoning
- When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
- An Empirical Study on the Effectiveness of Large Language Models for Binary Code Understanding
- MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model Framework
- MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
- Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
- General Preference Reinforcement Learning
- Learning from Language Feedback via Variational Policy Distillation
- MetaMoE: Diversity-Aware Proxy Selection for Privacy-Preserving Mixture-of-Experts Unification
- Real-Time Group Dynamics with LLM Facilitation: Evidence from a Charity Allocation Task
- Synthetic Interaction Data for Scalable Personalization in Large Language Models
- Targeted Neuron Modulation via Contrastive Pair Search
- Fair outputs, Biased Internals: Causal Potency and Asymmetry of Latent Bias in LLMs for High-Stakes Decisions
- Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation
- A Mechanistic Investigation of Supervised Fine Tuning
- Embeddings for Preferences, Not Semantics
- OmicsLM: A Multimodal Large Language Model for Multi-Sample Omics Reasoning
- Understanding Annotator Safety Policy with Interpretability
- SoK: Robustness in Large Language Models against Jailbreak Attacks
- Strat-Reasoner: Reinforcing Strategic Reasoning of LLMs in Multi-Agent Games
- Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
- On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
- Adjust to reality: LLM-driven test-time semantic adjustment for zero-shot fault diagnosis
- Fostering Self-Directed Growth with Generative AI: Toward a New Learning Analytics Framework
- Token-Efficient RL for LLM Reasoning
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
- Detecting Manipulated Contents Using Knowledge-Grounded Inference
- NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models
- Toward Efficient Exploration by Large Language Model Agents
- Multimodal Large Language Models for Medicine: A Comprehensive Survey
- Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models
- Prompt Injection Attack to Tool Selection in LLM Agents
- Model-based controller assisted domain randomization in deep reinforcement learning: application to nonlinear powertrain control
- PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement
- Contagion Networks: Evaluator Preference Propagation in Multi-Agent LLM Systems
- Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection
- On the Geometry of On-Policy Distillation
- Measuring and Mitigating Persona Distortions from AI Writing Assistance
- Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents
- Ethics Testing: Proactive Identification of Generative AI System Harms
- HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models
- Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling
- OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration
- Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
- TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation
- Reinforcing Human Behavior Simulation via Verbal Feedback
- Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance
- Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
- TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training
- Sequential Data Poisoning in LLM Post-Training
- MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents
- Learning is Forgetting: LLM Training As Lossy Compression
- Computational Hermeneutics: Evaluating generative AI as a cultural technology
- Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models
- MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
- Delightful Policy Gradient
- Matching Features, Not Tokens: Energy-Based Fine-Tuning of Language Models
- EEG-Based Brain-LLM Interface for Human Preference Aligned Generation
- The Tool-Overuse Illusion: Why Does LLM Prefer External Tools over Internal Knowledge?
- Robust AI Evaluation through Maximal Lotteries
- Buy versus Build an LLM: A Decision Framework for Governments
- Weak-Driven Learning: How Weak Agents make Strong Agents Stronger
- Rethinking the Trust Region in LLM Reinforcement Learning
- Maximum Likelihood Reinforcement Learning
- Feedback by Design: Understanding and Overcoming User Feedback Barriers in Conversational Agents
- Minerva: Reinforcement Learning with Verifiable Rewards for Cyber Threat Intelligence LLMs
- A Component-Based Survey of Interactions between Large Language Models and Multi-Armed Bandits
- Mastering the Game of Go with Self-play Experience Replay
- semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
- JailbreaksOverTime: Detecting Jailbreak Attacks Under Distribution Shift
- Modular Machine Learning: An Indispensable Path towards New-Generation Large Language Models
- Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
- Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI
- Conflicts in Texts: Data, Implications and Challenges
- GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
- GenCLS++: Pushing the Boundaries of Generative Classification in LLMs Through Comprehensive SFT and RL Studies Across Diverse Datasets
- Contextual Online Uncertainty-Aware Preference Learning for Human Feedback
- Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
- Beyond Static Snapshots: A Grounded Evaluation Framework for Language Models at the Agentic Frontier
- DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
- SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
- The State of AI Governance Research: AI Safety and Reliability in Real World Commercial Deployment
- Anyprefer: An Agentic Framework for Preference Data Synthesis
- KETCHUP: K-Step Return Estimation for Sequential Knowledge Distillation
- Graph of Attacks: Improved Black-Box and Interpretable Jailbreaks for LLMs
- Seeing the Goal, Missing the Truth: Human Accountability for AI Bias
- Meta-Learning in Self-Play Regret Minimization
- TLoRA: Tri-Matrix Low-Rank Adaptation of Large Language Models
- Investigating Co-Constructive Behavior of Large Language Models in Explanation Dialogues
- Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection
- RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
- Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
- Reason Like a Radiologist: Chain-of-Thought and Reinforcement Learning for Verifiable Report Generation
- Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
- TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs
- Narrative over Numbers: The Identifiable Victim Effect and its Amplification Under Alignment and Reasoning in Large Language Models
- The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
- Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing
- SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
- Think in Sentences: Explicit Sentence Boundaries Enhance Language Model's Capabilities
- Patches of Nonlinearity: Instruction Vectors in Large Language Models
- When Agents Say One Thing and Do Another: Validating Elicited Beliefs from LLMs
- The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
- Designing Digital Humans with Ambient Intelligence
- Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
- Debiasing LLMs by Fine-tuning
- The Spectral Geometry of Thought: Phase Transitions, Instruction Reversal, Token-Level Dynamics, and Perfect Correctness Prediction in How Transformers Reason
- Beyond Semantic Manipulation: Token-Space Attacks on Reward Models
- Reinforced Attention Learning
- Evaluating ChatGPT on Medical Information Extraction Tasks: Performance, Explainability and Beyond
- RPO-RAG: Aligning Small LLMs with Relation-aware Preference Optimization for Knowledge Graph Question Answering
- Exploring the Use of VLMs for Navigation Assistance for People with Blindness and Low Vision
- Large Language Models for Departmental Expert Review Quality Scores
- Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
- GameTalk: Training LLMs for Strategic Conversation
- Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
- SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding
- SolarGPT-QA: A Domain-Adaptive Large Language Model for Educational Question Answering in Space Weather and Heliophysics
- Can Instructed Retrieval Models Really Support Exploration?
- SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation
- When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
- Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning
- Whose Facts Win? LLM Source Preferences under Knowledge Conflicts
- MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory
- SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs
- LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
- Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems
- Do Chatbot LLMs Talk Too Much? The YapBench Benchmark
- LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
- Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili
- Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection
- Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training
- Accountability Asymmetry and Structural Trust in Autonomous AI Systems
- Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory
- Unified Lookup-Table Inference with Signed-Digit K/V Caches for Ternary LLMs
- Super Co-alignment of Human and AI for Sustainable Symbiotic Society
- ROAD: Reflective Optimization via Automated Debugging for Zero-Shot Agent Alignment
- Evaluating Parameter Efficient Methods for RLVR
- The Violation State: Safety State Persistence in a Multimodal Language Model Interface
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- Addressing Concept Mislabeling in Concept Bottleneck Models Through Preference Optimization
- Spark: A System for Scientifically Creative Idea Generation
- AI Awareness
- SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling
- Rubrics as Privileged Information for Open-Ended Generation
- Safety in Large Reasoning Models: A Survey
- Training Large Language Models to Reason via EM Policy Gradient
- JITServe: SLO-aware LLM Serving with Imprecise Request Information
- Generative Optimization for Incentivized Advertising with Global Level Constraints
- CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models
- Emergence of Reputation-Based Cooperation in LLM Agents
- GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction
- "Allow" to Achieve, Over-Privileged Inadvertently: The Unintended Cost of Task-Completion-Driven Pop-up Decisions in Mobile GUI Agents
- A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models
- Quantization Effects on Biomedical LLM Reliability
- Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning
- Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning
- Strengthening Target-Language Features: SAE-Based Steering for Multilingual Inference
- Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent
- Private Direct Preference Optimization for LLM Alignment
- Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems
- FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables
- Collab-RAG: Boosting Retrieval-Augmented Generation for Complex Question Answering via White-Box and Black-Box LLM Collaboration
- Self-Adaptive Cognitive Debiasing for Large Language Models in Decision-Making
- Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
- T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models
- Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long Contexts
- Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use
- FISH-Tuning: Enhancing PEFT Methods with Fisher Information
- Can ChatGPT Learn My Life From a Week of First-Person Video?
- Ensuring Safety in an Uncertain Environment: Constrained MDPs via Stochastic Thresholds
- The Curse of CoT: On the Limitations of Chain-of-Thought in In-Context Learning
- Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training
- Exploring ChatGPT and its impact on society
- AffectAgent: Collaborative Multi-Agent Reasoning for Retrieval-Augmented Multimodal Emotion Recognition
- Document Optimization for Black-Box Retrieval via Reinforcement Learning
- Can Post-Training Transform LLMs into Causal Reasoners?
- Learning to Diagnose and Correct Errors: Towards Moral Sensitivity Acquisition in Large Language Models
- From Brute Force to Semantic Insight: Performance-Guided Data Transformation Design with LLMs
- ParetoHqD: Fast Offline Multiobjective Alignment of Large Language Models using Pareto High-quality Data
- Enhancing LLM-Based Agents via Global Planning and Hierarchical Execution
- Taxonomy-Aware Evaluation of Vision-Language Models
- Safety Pretraining: Toward the Next Generation of Safe AI
- A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
- Steering the CensorShip: Uncovering Representation Vectors for LLM "Thought" Control
- Algorithm Discovery With LLMs: Evolutionary Search Meets Reinforcement Learning
- GreenMind: A Next-Generation Vietnamese Large Language Model for Structured and Logical Reasoning
- PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression
- Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
- Planning with Diffusion Models for Target-Oriented Dialogue Systems
- Private Federated Learning using Preference-Optimized Synthetic Data
- A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms
- Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion
- In which fields do ChatGPT 4o scores align better than citations with research quality?
- ZeroED: Hybrid Zero-shot Error Detection through Large Language Model Reasoning
- AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
- Kill two birds with one stone: generalized and robust AI-generated text detection via dynamic perturbations
- TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
- The Bitter Lesson Learned from 2,000+ Multilingual Benchmarks
- LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
- Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction
- TTRL: Test-Time Reinforcement Learning
- Fusing Reward and Dueling Feedback in Stochastic Bandits
- Dynamic Early Exit in Reasoning Models
- WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents
- SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
- Stop Summation: Min-Form Credit Assignment Is All Process Reward Model Needs for Reasoning
- Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
- MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
- Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
- Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs
- DSPO: Direct Semantic Preference Optimization for Real-World Image Super-Resolution
- Safety Co-Option and Compromised National Security: The Self-Fulfilling Prophecy of Weakened AI Risk Thresholds
- Establishing Reliability Metrics for Reward Models in Large Language Models
- Think2SQL: Reinforce LLM Reasoning Capabilities for Text2SQL
- Acting Less is Reasoning More! Teaching Model to Act Efficiently
- MARFT: Multi-Agent Reinforcement Fine-Tuning
- Contemplative Agent
- In-context Ranking Preference Optimization
- PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines
- Reinforcement Learning from Multi-level and Episodic Human Feedback
- Pairwise or Pointwise? Evaluating Feedback Protocols for Bias in LLM-Based Evaluation
- An LLM-enabled Multi-Agent Autonomous Mechatronics Design Framework
- A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
- Cross-Asset Risk Management: Integrating LLMs for Real-Time Monitoring of Equity, Fixed Income, and Currency Markets
- Reliable Confidence Intervals for Information Retrieval Evaluation Using Generative A.I.
- SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization
- Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
- FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
- LoRe: Personalizing LLMs via Low-Rank Reward Modeling
- Code2API: A Tool for Generating Reusable APIs from Stack Overflow Code Snippets
- Bias Analysis and Mitigation through Protected Attribute Detection and Regard Classification
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Science Hierarchography: Hierarchical Organization of Science Literature
- Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
- Analysing the Robustness of Vision-Language-Models to Common Corruptions
- Do Prompt Patterns Affect Code Quality? A First Empirical Assessment of ChatGPT-Generated Code
- Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling
- Prejudge-Before-Think: Enhancing Large Language Models at Test-Time by Process Prejudge Reasoning
- CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation
- Evaluating large language models on a highly-specialized topic, radiation oncology physics
- ChatGPT Empowered Long-Step Robot Control in Various Environments: A Case Application
- SMPL-GPTexture: Dual-View 3D Human Texture Estimation using Text-to-Image Generation Models
- Governance Challenges in Reinforcement Learning from Human Feedback: Evaluator Rationality and Reinforcement Stability
- Aligning Constraint Generation with Design Intent in Parametric CAD
- Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo
- Energy-Based Reward Models for Robust Language Model Alignment
- GraphAttack: Exploiting Representational Blindspots in LLM Safety Mechanisms
- Design Topological Materials by Reinforcement Fine-Tuned Generative Model
- VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization
- NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
- Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment
- Dynamic Hedging Strategies in Derivatives Markets with LLM-Driven Sentiment and News Analytics
- A Benchmark for End-to-End Zero-Shot Biomedical Relation Extraction with LLMs: Experiments with OpenAI Models
- MAIN: Mutual Alignment Is Necessary for instruction tuning
- Data-efficient LLM Fine-tuning for Code Generation
- SPIE: Semantic and Structural Post-Training of Image Editing Diffusion Models with AI feedback
- JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration
- AdaCoder: An Adaptive Planning and Multi-Agent Framework for Function-Level Code Generation
- Multilingual Contextualization of Large Language Models for Document-Level Machine Translation
- Collaboration and Controversy Among Experts: Rumor Early Detection by Tuning a Comment Generator
- Unsupervised Classification of English Words Based on Phonological Information: Discovery of Germanic and Latinate Clusters
- Window Token Concatenation for Efficient Visual Large Language Models
- MSL: Not All Tokens Are What You Need for Tuning LLM as a Recommender
- d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning
- An LLM-as-a-judge Approach for Scalable Gender-Neutral Translation Evaluation
- OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
- Can Pre-training Indicators Reliably Predict Fine-tuning Outcomes of LLMs?
- Evaluating the Diversity and Quality of LLM Generated Content
- Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics
- ChronoVision: Temporal Reasoning via Latent State Reconstruction
- From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs
- Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs
- Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration
- TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure
- MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration
- Cautious Context Steering for Language Model Personalization
- Cleo: A Transparent and Controllable Chatbot for Conversational Commerce
- Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation
- Position: It's Time to Optimize LLMs for Self-Consistency
- PolyAlign: Conditional Human-Distribution Alignment
- Persona-Pruner: Sculpting Lightweight Models for Role-Playing
- Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
- Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models
- SemiAdapt-Instruct: Extensible Instruction Tuning via Latent Domain-Specialised Adapters
- Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision Support
- RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs
- Memory in the LLM Era: Modular Architectures and Strategies in a Unified Framework
- Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning
- MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
- All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Generalizable Training
- Efficient Reasoning Models: A Survey
- Diffusion Distillation With Direct Preference Optimization For Efficient 3D LiDAR Scene Completion
- A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
- Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
- Agent-Q: Fine-Tuning Large Language Models for Quantum Circuit Generation and Optimization
- ReZero: Enhancing LLM search ability by trying one-more-time
- REWARD CONSISTENCY: Improving Multi-Objective Alignment from a Data-Centric Perspective
- Offline Learning and Forgetting for Reasoning with Large Language Models
- Reinforcing Compositional Retrieval: Retrieving Step-by-Step for Composing Informative Contexts
- How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
- InstructEngine: Instruction-driven Text-to-Image Alignment
- Augmented Relevance Datasets with Fine-Tuned Small LLMs
- Reasoning without Regret
- Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
- Summarizing Online Patient Conversations Using Generative Language Models: Experimental and Comparative Study
- Refining Financial Consumer Complaints through Multi-Scale Model Interaction
- Better Estimation of the Kullback--Leibler Divergence Between Language Models
- Joint Action Language Modelling for Transparent Policy Execution
- The Human Visual System Can Inspire New Interaction Paradigms for LLMs
- OctGPT: Octree-based Multiscale Autoregressive Models for 3D Shape Generation
- Enhancing Reasoning Abilities of Small LLMs with Cognitive Alignment
- Learning from Reference Answers: Versatile Language Model Alignment without Binary Human Preference Data
- DICE: A Framework for Dimensional and Contextual Evaluation of Language Models
- CHARM: Calibrating Reward Models With Chatbot Arena Scores
- Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
- Improving In-Context Learning with Reasoning Distillation
- RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability
- Fine-tuning a Large Language Model for Automating Computational Fluid Dynamics Simulations
- ControlNET: A Firewall for RAG-based LLM System
- Enhancing Mathematical Reasoning in Large Language Models with Self-Consistency-Based Hallucination Detection
- AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
- QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model
- SaRO: Enhancing LLM Safety through Reasoning-based Alignment
- Integrating Large Language Models for Automated Structural Analysis
- Alleviating the Fear of Losing Alignment in LLM Fine-tuning
- Continuum-Interaction-Driven Intelligence: Human-Aligned Neural Architecture via Crystallized Reasoning and Fluid Generation
- PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks
- Feature-Aware Malicious Output Detection and Mitigation
- On The Landscape of Spoken Language Models: A Comprehensive Survey
- Large Language Models Could Be Rote Learners
- X-Guard: Multilingual Guard Agent for Content Moderation
- Playpen: An Environment for Exploring Learning Through Conversational Interaction
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- Plan-and-Refine: Diverse and Comprehensive Retrieval-Augmented Generation
- NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-time Computation
- Enhancing Player Enjoyment with a Two-Tier DRL and LLM-Based Agent System for Fighting Games
- LoRI: Reducing Cross-Task Interference in Multi-Task Low-Rank Adaptation
- Toward Holistic Evaluation of Recommender Systems Powered by Generative Models
- CAReDiO: Cultural Alignment via Representativeness and Distinctiveness Guided Data Optimization
- EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
- Societal Impacts Research Requires Benchmarks for Creative Composition Tasks
- PinRec: Unified Generative Retrieval for Pinterest Recommender Systems
- From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
- Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
- On the Suitability of Reinforcement Fine-Tuning to Visual Tasks
- Stratified Expert Cloning for Retention-Aware Recommendation at Scale
- FactGuard: Leveraging Multi-Agent Systems to Generate Answerable and Unanswerable Questions for Enhanced Long-Context LLM Extraction
- Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
- InstructMPC: A Human-LLM-in-the-Loop Framework for Context-Aware Control
- Separator Injection Attack: Uncovering Dialogue Biases in Large Language Models Caused by Role Separators
- Multimedia and Visual Analytics in the Agentic Era
- Adversarial Training of Reward Models
- Llama-3-Nanda-10B-Chat: An Open Generative Large Language Model for Hindi
- COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values
- User Feedback Alignment for LLM-powered Exploration in Large-scale Recommendation Systems
- Lightweight and Direct Document Relevance Optimization for Generative Information Retrieval
- The Human Robot Social Interaction (HSRI) Dataset: Benchmarking Foundational Models' Social Reasoning
- SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
- CARE: Multilingual Human Preference Learning for Cultural Awareness
- Fast Controlled Generation from Language Models with Adaptive Weighted Rejection Sampling
- Generative Large Language Model usage in Smart Contract Vulnerability Detection
- AI alignment [wikipedia]
- AI safety [wikipedia]
- ChatGPT [wikipedia]
- Post-training of large language models [wikipedia]
- Products and applications of OpenAI [wikipedia]
- Reasoning model [wikipedia]
- Reinforcement learning [wikipedia]
- Reinforcement learning from human feedback [wikipedia]
- Paul Christiano [wikipedia]
Discussions
- Gemini lies to user about health info, says it wanted to make him feel better— Though commonly reported, Google doesn't consider it a security problem when models make things up [lemmy, 104 points, 21 comments]
- "Our models are neither fully aligned nor fully safe; they still generate toxic or biased outputs, make up facts, and generate sexual and violent content without explicit prompting." arxiv.org/abs/22 [bsky, 12 points, 1 comments]
- Training language models to follow instructions with human feedback [pdf] [hn, 3 points, 0 comments]
- Training language models to follow instructions with human feedback [hn, 2 points, 0 comments]
- arxiv.org/abs/2203.02155 Instruct GPT [bsky, 0 points, 0 comments]
- Possibly it does, but only because ChatGPT has in its pedigree InstructGPT, which was specifically trained w/ RLHF on prompts that contained instructions to do things? arxiv.org/abs/2203.02155 [bsky, 0 points, 0 comments]
- 15. InstructGPT arxiv.org/abs/2203.02155 Reading papers is one thing. Understanding why each changed the field is what makes you a stronger AI engineer. [bsky, 0 points, 0 comments]
- ... zu InstructGPT (arxiv.org/abs/2203.02155) mit der LLM nach dem ersten Pre-Training das Frage-/Antwort-Format gelernt wird. Fun Fakt: Einige Crowd Sourced Trainingsdaten versuchten dem erst pre-tra [bsky, 0 points, 1 comments]
- 人間からの入力をフィードバックにする手法はここらへん arxiv.org/abs/2203.02155 > Starting with a set of labeler-written prompts and prompts submitted through the OpenAI API, we collect a dataset of labeler demonstrations of th [bsky, 0 points, 0 comments]
Related