Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
2023/12/14 by Collin Burns, Pavel Izmailov, Burns, Collin +22 · 2 voices · 103 citations
Computer Science · #Computation and Language (cs.CL) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2312.09390
openalex publication_date 2023/12/14 · arxiv published 2023/12/14 · arxiv updated 2023/12/14 · openalex created_date 2023/12/19 · openalex updated_date 2026/07/28
Abstract
Widely used alignment techniques, such as reinforcement learning from human feedback (RLHF), rely on the ability of humans to supervise model behavior - for example, to evaluate whether a model faithfully followed instructions or generated safe outputs. However, future superhuman models will behave in complex ways too difficult for humans to reliably evaluate; humans will only be able to weakly supervise superhuman models. We study an analogy to this problem: can weak model supervision elicit the full capabilities of a much stronger model? We test this using a range of pretrained language models in the GPT-4 family on natural language processing (NLP), chess, and reward modeling tasks. We find that when we naively finetune strong pretrained models on labels generated by a weak model, they consistently perform better than their weak supervisors, a phenomenon we call weak-to-strong generalization. However, we are still far from recovering the full capabilities of strong models with naive finetuning alone, suggesting that techniques like RLHF may scale poorly to superhuman models without further work. We find that simple methods can often significantly improve weak-to-strong generalization: for example, when finetuning GPT-4 with a GPT-2-level supervisor and an auxiliary confidence loss, we can recover close to GPT-3.5-level performance on NLP tasks. Our results suggest that it is feasible to make empirical progress today on a fundamental challenge of aligning superhuman models.
Cited by
- Automated alignment is harder than you think
- When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops
- SERM: Self-Evolving Relevance Model with Agent-Driven Learning from Massive Query Streams
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization
- Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
- Uncertainty-Aware Data-Efficient AI: An Information-Theoretic Perspective
- A Flexible Multi-Agent LLM-Human Framework for Fast Human Validated Tool Building
- AgentShield: Make MAS more secure and efficient
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- Aligning Artificial Superintelligence via a Multi-Box Protocol
- Exploring Weak-to-Strong Generalization for CLIP-based Classification
- Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
- The Oversight Game: Learning to Cooperatively Balance an AI Agent's Safety and Autonomy
- Weak-to-Strong On-Policy Distillation
- Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
- Towards Scalable Oversight via Partitioned Human Supervision
- Weak-to-Strong Generalization under Distribution Shifts
- Position: Require Frontier AI Labs To Release Small "Analog" Models
- Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment
- Tandem Training for Language Models
- HoneyBee: Data Recipes for Vision-Language Reasoners
- AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?
- LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
- Users as Annotators: LLM Preference Learning from Comparison Mode
- Contrastive Weak-to-strong Generalization
- Mutual Learning for Hashing: Unlocking Strong Hash Functions from Weak Supervision
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards
- Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
- Does Weak-to-strong Generalization Happen under Spurious Correlations?
- On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
- Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration
- MORPH: PDE Foundation Models with Arbitrary Data Modality
- DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes
- WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
- Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
- Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
- Limitations of refinement methods for weak to strong generalization
- ACE and Diverse Generalization via Selective Disagreement
- Large-Small Model Collaborative Framework for Federated Continual Learning
- Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
- Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
- SKATE, a Scalable Tournament Eval: Weaker LLMs differentiate between stronger ones using verifiable challenges
- Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
- Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
- ε-Softmax: Approximating One-Hot Vectors for Mitigating Label Noise
- Lessons from complex systems science for AI governance
- Against racing to AGI: Cooperation, deterrence, and catastrophic risks
- Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
- Shaping capabilities with token-level data filtering
- SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
- TableDreamer: Progressive and Weakness-guided Data Synthesis from Scratch for Table Instruction Tuning
- Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
- The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
- Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients
- Training Language Model to Critique for Better Refinement
- Persona Features Control Emergent Misalignment
- Is Long-to-Short a Free Lunch? Investigating Inconsistency and Reasoning Efficiency in LRMs
- DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning
- Control Tax: The Price of Keeping AI in Check
- Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment
- Conformal Arbitrage: Risk-Controlled Balancing of Competing Objectives in Language Models
- On the Emergence of Weak-to-Strong Generalization: A Bias-Variance Perspective
- Cascading Adversarial Bias from Injection to Distillation in Language Models
- BIRD: Behavior Induction via Representation-structure Distillation
- Generalizable Video Quality Assessment via Weak-to-Strong Learning
- EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
- Learning to Reason without External Rewards
- Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplifications and Resistance in Multi-Agent Based LLM-as-Judge
- Continuous Self-Improvement of Large Language Models by Test-time Training with Verifier-Driven Sample Selection
- Incentivizing Strong Reasoning from Weak Supervision
- Calibrating Pre-trained Language Classifiers on LLM-generated Noisy Labels via Iterative Refinement
- Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer
- On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
- Value-Guided Search for Efficient Chain-of-Thought Reasoning
- Joint Flashback Adaptation for Forgetting-Resistant Instruction Tuning
- PEER pressure: Model-to-Model Regularization for Single Source Domain Generalization
- Denoising Mutual Knowledge Distillation in Bi-Directional Multiple Instance Learning
- T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
- Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
- DeepCritic: Deliberate Critique with Large Language Models
- R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
- Think First, Diffuse Fast: Improving Diffusion Language Model Reasoning via Autoregressive Plan Conditioning
- Weak-to-Strong Generalization via Direct On-Policy Distillation
- Calibrating Conservatism for Scalable Oversight
- Weak-Driven Learning: How Weak Agents make Strong Agents Stronger
- Federation over Text: Insight Sharing for Multi-Agent Reasoning
- Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks
- Super Co-alignment of Human and AI for Sustainable Symbiotic Society
- Scaling Laws For Scalable Oversight
- The Rise of Small Language Models in Healthcare: A Comprehensive Survey
- Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion
- Synergistic Weak-Strong Collaboration by Aligning Preferences
- Let Me Grok for You: Accelerating Grokking via Embedding Transfer from a Weaker Model
- Collaborative Learning of On-Device Small Model and Cloud-Based Large Model: Advances and Future Directions
- Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models
- Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
- Ctrl-Z: Controlling AI Agents via Resampling
- Dynamic Residual Safe Reinforcement Learning for Multi-Agent Safety-Critical Scenarios Decision-Making
- Mechanistic Anomaly Detection for "Quirky" Language Models
- Alice: Proactive Learning with Teacher's Demonstrations for Weak-to-Strong Generalization
- Weak-for-Strong: Training Weak Meta-Agent to Harness Strong Executors
Discussions
Related