RANK ANALYSIS OF INCOMPLETE BLOCK DESIGNS
1952/01/01 by RALPH ALLAN BRADLEY, Ralph A. Bradley, Milton E. Terry +1 · 1,382 citations
Economics, Econometrics and Finance · Engineering · Mathematics · #Block (permutation group theory) #Combinatorics #Credit Risk and Financial Regulations #Mathematics #Modeling, Simulation, and Optimization #Rank (graph theory) #Statistics #VLSI and FPGA Design Techniques
paper · doi:10.1093/biomet/39.3-4.324
published in Biometrika 39(3-4), 324-345 (Oxford University Press)
openalex publication_date 1952/01/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/02
Cited by
- Perceiving Protest: How Publics View the Disruptiveness and Effectiveness of Protest
- Convergence analysis of a family of Zermelo-type iterations for the Bradley--Terry model
- RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
- Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
- When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
- Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels
- Rater State Bias in RLHF Preference Data: An Audit Framework
- OR Else: A Differentiable Trust Region for Policy Optimization
- Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation
- Rationalizing Boltzmann Rationality: An Axiomatic Characterization of Entropy-Regularized Policies
- A Method for Learning Value Systems in Generative AI
- When Can You Debias an LLM Judge? Identifiability Limits, a Test, and Designs for Top-k Ranking
- TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue
- Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
- SODA: Semi On-Policy Black-Box Distillation for Large Language Models
- Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function
- Multi-Turn On-Policy Distillation with Prefix Replay
- Risk-Aware Preference Learning for Stochastic Outcomes
- Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
- From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
- Robust Critics: Defending LLMs Against Multi-Turn Attacks
- When AI Takes Sides on Questions of Faith: Persistent Asymmetries in AI-Mediated Faith Guidance
- Large-scale online deanonymization with LLMs
- Measuring Scalar Constructs in Social Science with LLMs
- ProxAnn: Use-Oriented Evaluations of Topic Models and Document Clustering
- LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
- Robin: A multi-agent system for automating scientific discovery
- Base Models Beat Aligned Models at Randomness and Creativity
- Explaining Human Preferences via Metrics for Structured 3D Reconstruction
- Decoding-based Regression
- Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings
- Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
- Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
- SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
- What do Reward Models Memorize?
- GEMCo: A Validated, Ethically Releasable Proxy for Inaccessible Counselling Data
- Simulating Single Transferable Voting for the Colorado House of Representatives
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
- Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining
- VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
- Group Preference Collapse in Personalized Multimodal Large Language Models
- Mean-Field Analysis and Optimal Control of a Dynamic Rating and Matchmaking System
- SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams
- Statistical and computational challenges in ranking
- Safety Alignment of LMs via Non-cooperative Games
- Counterfactual LLM-based Framework for Measuring Rhetorical Style
- From Pixels to Predicates Structuring urban perception with scene graphs
- Efficient Personalization of Generative Models via Optimal Experimental Design
- Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
- Model inference for ranking from pairwise comparisons
- Vibe Spaces for Creatively Connecting and Expressing Visual Concepts
- DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization
- Estimating problem difficulty without ground truth using Large Language Model comparisons
- Explainable reinforcement learning from human feedback to improve alignment
- ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
- Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
- From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label Classification
- LLMs Can Assist with Proposal Selection at Large User Facilities
- RoleRMBench & RoleRM: Towards Reward Modeling for Profile-Based Role Play in Dialogue Systems
- Multi-dimensional Preference Alignment by Conditioning Reward Itself
- T-pro 2.0: An Efficient Russian Hybrid-Reasoning Model and Playground
- Targeting Misalignment: A Conflict-Aware Framework for Reward-Model-based LLM Alignment
- AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
- Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE
- Statistical analysis of Inverse Entropy-regularized Reinforcement Learning
- Learning Preferences and User Engagement Using Choice and Time Data
- When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
- DEAR: Dataset for Evaluating the Aesthetics of Rendering
- SA-IQA: Redefining Image Quality Assessment for Spatial Aesthetics with Multi-Dimensional Rewards
- Overcoming State Inertia: Minimally Invasive Temporal Alignment for Evolving Contexts
- Joint Progression Modeling (JPM): A Probabilistic Framework for Mixed-Pathology Progression
- Hypothesis Testing for Generalized Thurstone Models
- Distillability of LLM Security Logic: Predicting Attack Success Rate of Outline Filling Attack via Ranking Regression
- UMM-RM: An Upcycle-and-Merge MoE Reward Model for Mitigating Reward Hacking
- When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
- RecruitView: A Multimodal Dataset for Predicting Personality and Interview Performance for Human Resources Applications
- A Trainable Centrality Framework for Modern Data
- FLAWS: A Benchmark for Error Identification and Localization in Scientific Papers
- Real-Time Long Horizon Air Quality Forecasting via Group-Relative Policy Optimization
- Video Generation Models Are Good Latent Reward Models
- GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs
- Test-Time Preference Optimization for Image Restoration
- SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- Aligning Generative Music AI with Human Preferences: Methods and Challenges
- The Alignment Game: A Theory of Long-Horizon Alignment Through Recursive Curation
- Exploring question answering: metric analysis and evaluation framework for enhanced interpretability
- Probing Preference Representations: A Multi-Dimensional Evaluation and Analysis Method for Reward Models
- Maximizing the efficiency of human feedback in AI alignment: a comparative analysis
- Mitigating Length Bias in RLHF through a Causal Lens
- MetaGDPO: Alleviating Catastrophic Forgetting with Metacognitive Knowledge through Group Direct Preference Optimization
- PIRA: Preference-Oriented Instruction-Tuned Reward Models with Dual Aggregation
- Moment estimation in paired comparison models with a growing number of subjects
- Black-Box On-Policy Distillation of Large Language Models
- AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment
- Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
- SafeMIL: Learning Offline Safe Imitation Policy from Non-Preferred Trajectories
- Test-driven Reinforcement Learning in Continuous Control
- SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
- HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection
- SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
- Characterizing AI Manipulation Risks in Brazilian YouTube Climate Discourse
- Ties in Paired-Comparison Experiments: A Generalization of the Bradley-Terry Model
- CPO: Condition Preference Optimization for Controllable Image Generation
- Towards Reliable Human Evaluations in Gesture Generation: Insights from a Community-Driven State-of-the-Art Benchmark
- Diffusion-SDPO: Safeguarded Direct Preference Optimization for Diffusion Models
- The Bradley-Terry Stochastic Block Model
- Speech-Based Prioritization for Schizophrenia Intervention
- From Measurement to Expertise: Empathetic Expert Adapters for Context-Based Empathy in Conversational AI Agents
- Human-AI Collaboration with Misaligned Preferences
- Automated Reward Design for Gran Turismo
- Toward Objective and Interpretable Prosody Evaluation in Text-to-Speech: A Linguistically Motivated Approach
- Deployable Vision-driven UAV River Navigation via Human-in-the-loop Preference Alignment
- ThoughtProbe: Classifier-Guided LLM Thought Space Exploration via Probing Representations
- MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
- Value Drifts: Tracing Value Alignment During LLM Post-Training
- Offline Clustering of Preference Learning with Active-data Augmentation
- One Model to Critique Them All: Rewarding Agentic Tool-Use via Efficient Reasoning
- E-Scores for (In)Correctness Assessment of Generative Model Outputs
- Wilks' theorems in some exponential random graph models
- Mediocrity is the key for LLM as a Judge Anchor Selection
- Default Bayes factors for ANOVA designs
- Reward Models are Metrics in a Trench Coat
- Estimation from Pairwise Comparisons: Sharp Minimax Bounds with Topology Dependence
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- Prior Distributions for the Bradley-Terry Model of Paired Comparisons
- Fast and Accurate Image Super-Resolution with Deep Laplacian Pyramid Networks
- Generative Bayesian Optimization: Generative Models as Acquisition Functions
- GReF: A Unified Generative Framework for Efficient Reranking via Ordered Multi-token Prediction
- Learning Video-Story Composition via Recurrent Neural Network
- Greedy Sampling Is Provably Efficient for RLHF
- PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
- Semi-Supervised Preference Optimization with Limited Feedback
- Fortytwo: Swarm Inference with Peer-Ranked Consensus
- GIFT: Group-relative Implicit Fine Tuning Integrates GRPO with DPO and UNA
- Protein Language Model Fitness Is a Matter of Preference
- The Burden of Interactive Alignment with Inconsistent Preferences
- Debiasing Reward Models by Representation Learning with Guarantees
- Lightweight Robust Direct Preference Optimization
- RobotArena ∞: Scalable Robot Benchmarking via Real-to-Sim Translation
- POPI: Personalizing LLMs via Optimized Natural Language Preference Inference
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- Learning "Partner-Aware" Collaborators in Multi-Party Collaboration
- Estimating the Error of Large Language Models at Pairwise Text Comparison
- Generalized Top-k Mallows Model for Ranked Choices
- Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
- Cost-Sensitive Freeze-thaw Bayesian Optimization for Efficient Hyperparameter Tuning
- Weak-to-Strong Generalization under Distribution Shifts
- Grouped sparse paired comparisons in the Bradley-Terry model
- Self-evolving expertise in complex non-verifiable subject domains: dialogue as implicit meta-RL
- Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
- Capturing Intransitive Dominance in Tennis Forecasting: A Graph Neural Network Approach
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- Rank-GRPO: Training LLM-based Conversational Recommender Systems with Reinforcement Learning
- Why DPO is a Misspecified Estimator and How to Fix It
- Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models
- Rectifying Shortcut Behaviors in Preference-based Reward Learning
- Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
- Pairwise Choice Markov Chains
- Ranking-based Preference Optimization for Diffusion Models from Implicit User Feedback
- ADPO: Anchored Direct Preference Optimization
- DP2O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution
- Cultural Alien Sampler: Open-ended art generation balancing originality and coherence
- Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
- Eliciting Truthful Feedback for Preference-Based Learning via the VCG Mechanism
- Dual-Weighted Reinforcement Learning for Generative Preference Modeling
- Voting with the Graph: Stable RLAIF via Topological Consistency Maximization
- GraphMind: Interactive Novelty Assessment System for Accelerating Scientific Discovery
- Budget-aware Test-time Scaling via Discriminative Verification
- Confidence as a Reward: Transforming LLMs into Reward Models
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Towards Understanding Valuable Preference Data for Large Language Model Alignment
- On the Role of Preference Variance in Preference Optimization
- Single Image Deraining: From Model-Based to Data-Driven and Beyond
- DocReward: A Document Reward Model for Structuring and Stylizing
- APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport
- Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
- Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval
- RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model
- Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey
- Efficient Bayesian Inference from Noisy Pairwise Comparisons
- Users as Annotators: LLM Preference Learning from Comparison Mode
- Score-Based Density Estimation from Pairwise Comparisons
- Active Model Selection for Large Language Models
- What Makes a Visualization Image Complex?
- Contrastive Weak-to-strong Generalization
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
- From Data to Rewards: a Bilevel Optimization Perspective on Maximum Likelihood Estimation
- Agent Bain vs. Agent McKinsey: A New Text-to-SQL Benchmark for the Business Domain
- Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
- Towards Better Optimization For Listwise Preference in Diffusion Models
- Optimal Stopping vs Best-of-N for Inference Time Optimization
- Model-free Rank Aggregation in the Presence of Rater Heterogeneity: A Maximum Score Approach
- Online Rubrics Elicitation from Pairwise Comparisons
- Reward Model Perspectives: Whose Opinions Do Reward Models Reward?
- Provably Convergent Primal-Dual DPO for Constrained LLM Alignment
- On the Role of Difficult Prompts in Self-Play Preference Optimization
- Provably Mitigating Corruption, Overoptimization, and Verbosity Simultaneously in Offline and Online RLHF/DPO Alignment
- Margin Adaptive DPO: Leveraging Reward Model for Granular Control in Preference Optimization
- From Noisy Traces to Stable Gradients: Bias-Variance Optimized Preference Optimization for Aligning Large Reasoning Models
- SONA: Learning Conditional, Unconditional, and Mismatching-Aware Discriminator
- Aligning Language Models with Clinical Expertise: DPO for Heart Failure Nursing Documentation in Critical Care
- BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
- Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
- LegalSim: Multi-Agent Simulation of Legal Systems for Discovering Procedural Exploits
- CVSM: Contrastive Vocal Similarity Modeling
- Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
- Hybrid-MST: A Hybrid Active Sampling Strategy for Pairwise Preference Aggregation
- From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
- AP2O-Coder: Adaptively Progressive Preference Optimization for Reducing Compilation and Runtime Errors in LLM-Generated Code
- Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information
- Judging with Confidence: Calibrating Autoraters to Preference Distributions
- GRPO-λ: Credit Assignment improves LLM Reasoning
- Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
- fev-bench: A Realistic Benchmark for Time Series Forecasting
- Alignment-Aware Decoding
- Fading to Grow: Growing Preference Ratios via Preference Fading Discrete Diffusion for Recommendation
- RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation
- RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
- Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
- Limited Preference Data? Learning Better Reward Model with Latent Space Synthesis
- Information Design With Large Language Models
- Can Molecular Foundation Models Know What They Don't Know? A Simple Remedy with Preference Optimization
- Generative Value Conflicts Reveal LLM Priorities
- OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment
- RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
- Meta-Router: Bridging Gold-standard and Preference-based Evaluations in Large Language Model Routing
- Interactive Groupwise Comparison for Reinforcement Learning from Human Feedback
- Preference-Based Dynamic Ranking Structure Recognition
- STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
- AI can help humans find common ground in democratic deliberation
- Tagging the Thought: Unlocking Personalization Reasoning via Reinforcement Learning
- General Exploratory Bonus for Optimistic Exploration in RLHF
- Multiplayer Nash Preference Optimization
- Adaptive Margin RLHF via Preference over Preferences
- Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
- Think Socially via Cognitive Reasoning
- Preference-Guided Learning for Sparse-Reward Multi-Agent Reinforcement Learning
- X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
- Failure Modes of Maximum Entropy RLHF
- Future Policy Aware Preference Learning for Mathematical Reasoning
- Embedding Domain Knowledge for Large Language Models via Reinforcement Learning from Augmented Generation
- MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation
- MedLLM: An Open Medical Language Model at the Sub-Billion Scale
- Implicit Safety Alignment from Crowd Preferences
- Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
- HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models
- Beyond Brightening Low-light Images
- Mathematicians’ Assessments of the Explanatory Value of Proofs
- Optimal Pairwise Comparison Procedures for Subjective Evaluation
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- Similarity Field Theory: A Mathematical Framework for Intelligence
- Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
- Dynamic Classifier-Free Diffusion Guidance via Online Feedback
- GPO: Learning from Critical Steps to Improve LLM Reasoning
- Agentic Aerial Cinematography: From Dialogue Cues to Cinematic Trajectories
- Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search
- siDPT: siRNA Efficacy Prediction via Debiased Preference-Pair Transformer
- Aligning Audio Captions with Human Preferences
- Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
- Force-Modulated Visual Policy for Robot-Assisted Dressing with Arm Motions
- zELO: ELO-inspired Training Method for Rerankers and Embedding Models
- Designing Rules to Pick a Rule: Aggregation by Consistency
- Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
- Pluralistic Off-policy Evaluation and Alignment
- What Matters in Data for DPO?
- ToolRM: Outcome Reward Models for Tool-Calling Large Language Models
- When Are Two RLHF Objectives the Same?
- Term2Note: Synthesising Differentially Private Clinical Notes from Medical Terms
- Decoding Alignment: A Critical Survey of LLM Development Initiatives through Value-setting and Data-centric Lens
- Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration
- Compass-v3: Scaling Domain-Specific LLMs for Multilingual E-Commerce in Southeast Asia
- Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling
- RewardDance: Reward Scaling in Visual Generation
- Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
- CM-Align: Consistency-based Multilingual Alignment for Large Language Models
- Towards Generalized Routing: Model and Agent Orchestration for Adaptive and Efficient Inference
- Verbalized Algorithms: Classical Algorithms are All You Need (Mostly)
- Continuous Audio Language Models
- Let's Roleplay: Examining LLM Alignment in Collaborative Dialogues
- Legislators’ sentiment analysis supervised by legislators
- Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment
- Connections between reinforcement learning with feedback,test-time scaling, and diffusion guidance: An anthology
- GeoArena: Evaluating Open-World Geographic Reasoning in Large Vision-Language Models
- Post Hoc Regression Refinement via Pairwise Rankings
- A Foundation Model for Chest X-ray Interpretation with Grounded Reasoning via Online Reinforcement Learning
- SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
- Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
- Adaptive Preference Optimization with Uncertainty-aware Utility Anchor
- IDEAlign: Comparing Large Language Models to Human Experts in Open-ended Interpretive Annotations
- Towards Reasoning for PDE Foundation Models: A Reward-Model-Driven Inference-Time-Scaling Algorithm
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
- Unifi3D: A Study on 3D Representations for Generation and Reconstruction in a Common Framework
- Benchmarking Large Language Models for Personalized Guidance in AI-Enhanced Learning
- GRAM-R2: Self-Training Generative Foundation Reward Models for Reward Reasoning
- MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
- RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
- ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- Vevo2: A Unified and Controllable Framework for Speech and Singing Voice Generation
- Brain-Inspired Deep Networks for Image Aesthetics Assessment
- OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
- Interestingness First Classifiers
- Spotlight Attention: Towards Efficient LLM Generation via Non-linear Hashing-based KV Cache Retrieval
- Optimization Transfer Using Surrogate Objective Functions
- Search-Based Credit Assignment for Offline Preference-Based Reinforcement Learning
- Learning from Preferences and Mixed Demonstrations in General Settings
- Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
- QuarkMed Medical Foundation Model Technical Report
- Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings
- Controlling Multimodal LLMs via Reward-guided Decoding
- Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
- Fusing Rewards and Preferences in Reinforcement Learning
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- Performance of GPT-5 Frontier Models in Ophthalmology Question Answering
- On Negative-aware Preference Optimization for Recommendation
- Interpretable Reward Model via Sparse Autoencoder
- Who Pays the RENT? Implications of Spatial Inequality for Prediction-Based Allocation Policies
- Personalized Recommendations via Active Utility-based Pairwise Sampling
- Street-Level AI: Are Large Language Models Ready for Real-World Judgments?
- WeChat-YATT: A Scalable, Simple, Efficient, and Production Ready Training Library
- \(X\)-evolve: Solution space evolution powered by large language models
- Learning to Align, Aligning to Learn: A Unified Approach for Self-Optimized Alignment
- Data Selection for LLM Alignment Using Fine-Grained Preferences
- Pref-GUIDE: Continual Policy Learning from Real-Time Human Feedback via Preference-Based Learning
- Latent Preference Bandits
- Posterior-GRPO: Rewarding Reasoning Processes in Code Generation
- Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
- FaST: Feature-aware Sampling and Tuning for Personalized Preference Alignment with Limited Data
- Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap
- HPSv3: Towards Wide-Spectrum Human Preference Score
- Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
- HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation
- VAGPO: Vision-augmented Asymmetric Group Preference Optimization for Graph Routing Problems
- BioDisco: Multi-agent hypothesis generation with dual-mode evidence, iterative feedback and temporal evaluation
- Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
- Comparison of Large Language Models for Deployment Requirements
- G-Core: A Simple, Scalable and Balanced RLHF Trainer
- Post-Training Large Language Models via Reinforcement Learning from Self-Feedback
- Libra: Assessing and Improving Reward Model by Learning to Think
- Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
- TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
- Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
- Music Arena: Live Evaluation for Text-to-Music
- TADT-CSA: Temporal Advantage Decision Transformer with Contrastive State Abstraction for Generative Recommendation
- MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
- Can AI Model the Complexities of Human Moral Decision-making? A Qualitative Study of Kidney Allocation Decisions
- PITA: Preference-Guided Inference-Time Alignment for LLM Post-Training
- Designing User-Centric Metrics for Evaluation of Counterfactual Explanations
- RobotValues: Evaluating Household Robots When Human Values Conflict
- MathDuels: Evaluating LLMs as Problem Posers and Solvers
- V1: Unifying Generation and Self-Verification for Parallel Reasoners
- Statistical and Algorithmic Foundations of Reinforcement Learning
- Preference-based Multi-Objective Reinforcement Learning
- Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks
- Food and Food-Odor Preferences in Dogs: A Pilot Study
- Making Language Model a Hierarchical Classifier
- Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
- PrefPalette: Personalized Preference Modeling with Latent Attributes
- Learning to summarize user information for personalized reinforcement learning from human feedback
- Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
- QuRe: Query-Relevant Retrieval through Hard Negative Sampling in Composed Image Retrieval
- EnlightenGAN: Deep Light Enhancement Without Paired Supervision
- ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
- On Monotonicity in AI Alignment
- GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO
- Optimal Differentially Private Ranking from Pairwise Comparisons
- Generalizing while preserving monotonicity in comparison-based preference learning models
- Mallows Model with Learned Distance Metrics: Sampling and Maximum Likelihood Estimation
- Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
- Principled Foundations for Preference Optimization
- Stable Preference Optimization: A Bilevel Approach to Catastrophic Preference Shift
- Not All Preferences are What You Need for Post-Training: Selective Alignment Strategy for Preference Optimization
- Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
- Why is Your Language Model a Poor Implicit Reward Model?
- Divergence Minimization Preference Optimization for Diffusion Model Alignment
- Reward Models Can Improve Themselves: Reward-Guided Adversarial Failure Mode Discovery for Robust Reward Modeling
- Interpretable Reward Modeling with Active Concept Bottlenecks
- Learning Latent Reasoning Traces for Scalar Reward Models End-to-End
- News Source Citing Patterns in AI Search Systems
- wd1: Weighted Policy Optimization for Reasoning in Diffusion Language Models
- Discrete Diffusion Trajectory Alignment via Stepwise Decomposition
- Pre-Trained Policy Discriminators are General Reward Models
- Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
- Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth
- Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity
- GenAI-Powered Inference
- Listwise Preference Alignment Optimization for Tail Item Recommendation
- Data Diversification Methods In Alignment Enhance Math Performance In LLMs
- MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model Pretraining
- Confidence and Stability of Global and Pairwise Scores in NLP Evaluation
- Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities
- Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
- Residual Reward Models for Preference-based Reinforcement Learning
- SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
- Linearly Decoding Refused Knowledge in Aligned Language Models
- FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes
- Generalist Reward Models: Found Inside Large Language Models
- The Hidden Link Between RLHF and Contrastive Learning
- Explicit Preference Optimization: No Need for an Implicit Reward Model
- Preference Completion: Large-scale Collaborative Ranking from Pairwise Comparisons
- Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
- Data-driven satisficing measure and ranking
- Benchmarking Music Generation Models and Metrics via Human Preference Studies
- TC-Light: Temporally Coherent Generative Rendering for Realistic World Transfer
- LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
- 3D Arena: An Open Platform for Generative 3D Evaluation
- The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical Proofs
- Shrinking the Generation-Verification Gap with Weak Verifiers
- RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
- Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basque
- Advancing Harmful Content Detection in Organizational Research: Integrating Large Language Models with Elo Rating System
- Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
- Reranking-based Generation for Unbiased Perspective Summarization
- Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
- User-Guided Force-Directed Graph Layout
- Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
- Reward Models in Deep Reinforcement Learning: A Survey
- Reward Model Interpretability via Optimal and Pessimal Tokens
- DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimization
- GRAM: A Generative Foundation Reward Model for Reward Generalization
- TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization
- Value-Free Policy Optimization via Reward Partitioning
- PB2: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning
- VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
- Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
- PersonaFeedback: A Large-scale Human-annotated Benchmark For Personalization
- AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models
- Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory
- From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment
- Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
- Personalized LLM Decoding via Contrasting Personal Preference
- Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference Optimization
- From Replication to Redesign: Exploring Pairwise Comparisons for LLM-Based Peer Review
- Rating competitors in games with strength-dependent tie probabilities
- Pareto Optimal Code Generation
- Large Language Models Can Be a Viable Substitute for Expert Political Surveys When a Shock Disrupts Traditional Measurement Approaches
- Debiasing Online Preference Learning via Preference Feature Preservation
- Preference Learning for AI Alignment: a Causal Perspective
- FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
- Reliable Evaluation of MRI Motion Correction: Dataset and Insights
- Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey
- SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
- Search Arena: Analyzing Search-Augmented LLMs
- RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation
- When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
- Identifying Terms and Conditions Important to Consumers using Crowdsourcing
- QQSUM: A Novel Task and Model of Quantitative Query-Focused Summarization for Review-based Product Question Answering
- BPO: Revisiting Preference Modeling in Direct Preference Optimization
- Policy-labeled Preference Learning: Is Preference Enough for RLHF?
- Robust Preference Optimization via Dynamic Target Margins
- Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation
- RewardAnything: Generalizable Principle-Following Reward Models
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
- TO-GATE: Clarifying Questions and Summarizing Responses with Trajectory Optimization for Eliciting Human Preference
- Understanding the Impact of Sampling Quality in Direct Preference Optimization
- Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences
- Beyond the Surface: Measuring Self-Preference in LLM Judgments
- Beyond Text Compression: Evaluating Tokenizers Across Scales
- Quantitative LLM Judges
- What Makes LLMs Effective Sequential Recommenders? A Study on Preference Intensity and Temporal Context
- SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis
- Rating Quality of Diverse Time Series Data by Meta-learning from LLM Judgment
- Assumption-free stability for ranking problems
- Beyond RLHF: A Unified Theoretical Framework of Alignment
- RewardBench 2: Advancing Reward Model Evaluation
- Doubly Robust Alignment for Large Language Models
- CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries
- K-order Ranking Preference Optimization for Large Language Models
- On Symmetric Losses for Robust Policy Optimization with Noisy Preferences
- Paired comparison models with strength-dependent ties and order effects
- MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning
- Evaluating Gemini in an arena for learning
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- Thompson Sampling in Online RLHF with General Function Approximation
- ChARM: Character-based Act-adaptive Reward Modeling for Advanced Role-Playing Language Agents
- REOrdering Patches Improves Vision Models
- Bayesian Optimization from Human Feedback: Near-Optimal Regret Bounds
- Learning Parametric Distributions from Samples and Preferences
- Bounded-Abstention Pairwise Learning to Rank
- Discriminative Policy Optimization for Token-Level Reward Models
- Towards Reward Fairness in RLHF: From a Resource Allocation Perspective
- LLM Agents for Bargaining with Utility-based Feedback
- Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
- Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
- Preference Learning with Response Time: Robust Losses and Guarantees
- Align-DA: Align Score-based Atmospheric Data Assimilation with Multiple Preferences
- LLMs Judging LLMs: A Simplex Perspective
- Reverse Preference Optimization for Complex Instruction Following
- Photography Perspective Composition: Towards Aesthetic Perspective Recommendation
- Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching
- SquareχPO: Differentially Private and Robust χ2-Preference Optimization in Offline Direct Alignment
- Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling
- Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback
- Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
- Risk-aware Direct Preference Optimization under Nested Risk Measure
- Energy-based Preference Optimization for Test-time Adaptation
- Preference Optimization by Estimating the Ratio of the Data Distribution
- Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries
- Learning a Pessimistic Reward Model in RLHF
- FairPO: Robust Preference Optimization for Fair Multi-Label Learning
- Marginal minimization and sup-norm expansions in perturbed optimization
- Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
- Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits
- Frictional Agent Alignment Framework: Slow Down and Don't Break Things
- Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
- SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
- A Survey on Progress in LLM Alignment from the Perspective of Reward Design
- Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
- LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models
- Incentivizing High-Quality Human Annotations with Golden Questions
- Reward Model Overoptimisation in Iterated RLHF
- Scalable Valuation of Human Feedback through Provably Robust Model Alignment
- Large Language Models Do Multi-Label Classification Differently
- Value-Guided Search for Efficient Chain-of-Thought Reasoning
- On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
- KL-regularization Itself is Differentially Private in Bandits and RLHF
- Shape it Up! Restoring LLM Safety during Finetuning
- MPO: Multilingual Safety Alignment via Reward Gap Optimization
- Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees
- Isotonic Bradley-Terry Model for Paired Comparison Data
- Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
- Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data?
- Direct Preference Optimization for Adaptive Concept-based Explanations
- Generalised Probabilistic Modelling and Improved Uncertainty Estimation in Comparative LLM-as-a-judge
- A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO
- SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models
- Imitation Learning via Focused Satisficing
- Unify Graph Learning with Text: Unleashing LLM Potentials for Session Search
- Preference Learning with Lie Detectors can Induce Honesty or Evasion
- Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency
- Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks
- Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
- Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
- Modeling Aesthetic Preferences in 3D Shapes: A Large-Scale Paired Comparison Study Across Object Categories
- Graph-Reward-SQL: Execution-Free Reinforcement Learning for Text-to-SQL via Graph Matching and Stepwise Reward
- Uncovering Competing Poisoning Attacks in Retrieval-Augmented Generation
- SGDPO: Self-Guided Direct Preference Optimization for Language Model Alignment
- Pairwise Calibrated Rewards for Pluralistic Alignment
- Online Iterative Self-Alignment for Radiology Report Generation
- AdaBoN: Adaptive Best-of-N Alignment
- Mutual-Taught for Co-adapting Policy and Reward Models
- J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
- Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
- REMOR: Automated Peer Review Generation with LLM Reasoning and Multi-Objective Reinforcement Learning
- A Systematic Analysis of Base Model Choice for Reward Modeling
- SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization
- Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
- WorldPM: Scaling Human Preference Modeling
- Maximum Selection and Sorting with Adversarial Comparators and an Application to Density Estimation
- Distributionally Robust Listwise Preference Optimization
- Bayesian and Motivated Reasoning in AI Agents
- Self-Consuming Generative Models with Adversarially Curated Data
- InfoPO: On Mutual Information Maximization for Large Language Model Alignment
- LCES: Zero-shot Automated Essay Scoring via Pairwise Comparisons Using Large Language Models
- On the Robustness of Reward Models for Language Model Alignment
- Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
- Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning
- Convert Language Model into a Value-based Strategic Planner
- Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes
- Inducing Robustness in a 2 Dimensional Direct Preference Optimization Paradigm
- Cannot See the Forest for the Trees: Invoking Heuristics and Biases to Elicit Irrational Choices of LLMs
- Learning Semantic Priorities for Autonomous Target Search
- Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
- Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation
- The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank
- JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation
- War of Words: The Competitive Dynamics of Legislative Processes
- War of Words II: Enriched Models of Law-Making Processes
- JP-TL-Bench: Anchored Pairwise LLM Evaluation for Bidirectional Japanese-English Translation
- Scorio.jl: A Julia package for ranking stochastic responses
- Transposition is Nearly Optimal for IID List Update
- Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
- Regularization in Paired Comparison Models via Pseudo-Games and Phantom Players
- Configurable Reward Model for Balanced Safety Alignment
- Causal methods for LLM development and evaluation
- Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
- JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment
- General Preference Reinforcement Learning
- A Hyperbolic Cosine Latent Trait Model For Unfolding Dichotomous Single-Stimulus Responses
- Do Large Language Models Always Tell The Same Stories?
- Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
- Why Global LLM Leaderboards Are Misleading: Small Portfolios for Heterogeneous Supervised ML
- Sequential Data Poisoning in LLM Post-Training
- Robust AI Evaluation through Maximal Lotteries
- GVPO: Group Variance Policy Optimization for Large Language Model Post-Training
- Calibrating Translation Decoding with Quality Estimation on LLMs
- Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
- Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
- Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks
- SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling
- A Dual Evaluation for Music Transcription
- Private Direct Preference Optimization for LLM Alignment
- FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables
- Advances in Human-Robot Handshaking
- Quantifying Response Dependence Between Two Dichotomous Items Using the Rasch Model
- Efficient Portfolio Selection through Preference Aggregation with Quicksort and the Bradley--Terry Model
- Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion
- From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
- CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial Optimization
- AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
- In-context Ranking Preference Optimization
- Reinforcement Learning from Multi-level and Episodic Human Feedback
- A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents
- Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning
- SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization
- LoRe: Personalizing LLMs via Low-Rank Reward Modeling
- Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling
- Energy-Based Reward Models for Robust Language Model Alignment
- SPIE: Semantic and Structural Post-Training of Image Editing Diffusion Models with AI feedback
- ADAPT: Actively Discovering and Adapting to Preferences for any Task
- FiSMiness: A Finite State Machine Based Paradigm for Emotional Support Conversations
- KVAE: Family of Tokenizers for Multimodal Generative Models
- Cautious Context Steering for Language Model Personalization
- GenGA: Editable and Data-Grounded Graphical Abstract Generation for Academic Papers
- Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation
- MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
- LazyReview A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews
- Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
- Discriminator-Free Direct Preference Optimization for Video Diffusion
- Network-based ranking in social systems: three challenges
- FuseRL: Dense Preference Optimization for Heterogeneous Model Fusion
- ThoughtProbe: Classifier-Guided Thought Space Exploration Leveraging LLM Intrinsic Reasoning
- Adversarial Training of Reward Models
- The uncanny valley effect in typically developing children and its absence in children with autism spectrum disorders. [europepmc]
- Scale Separation Reliability: What Does It Mean in the Context of Comparative Judgment? [europepmc]
- A Multidimensional IRT Approach for Dynamically Monitoring Ability Growth in Computerized Practice Environments. [europepmc]
- Tumor heterogeneity and acquired drug resistance in FGFR2-fusion-positive cholangiocarcinoma through rapid research autopsy. [europepmc]
- Geometric Affordance Perception: Leveraging Deep 3D Saliency With the Interaction Tensor. [europepmc]
- Attention-Guided Multi-Scale Feature Fusion Network for Low-Light Image Enhancement. [europepmc]
- Scaling preferences using probabilistic choice models: is there a ratio-scale representation of subjective liking? [europepmc]
- Value Propositions of Public Adult Hearing Rehabilitation in Denmark. [europepmc]
- Pic2Plate: A Vision-Language and Retrieval-Augmented Framework for Personalized Recipe Recommendations. [europepmc]
- Human-anchored longitudinal comparison of generative AI with a bias-calibrated LLM-as-judge. [europepmc]
- Ideometrics: a scientific approach to generating, evaluating, and prioritising ideas. [europepmc]
- Relationship Between Display Pixel Structure and Gloss Perception. [europepmc]