OpenAssistant Conversations -- Democratizing Large Language Model Alignment
2023/04/14 by Andreas Köpf, Yannic Kilcher, Köpf, Andreas +33 · 144 citations
Computer Science · #Topic Modeling #Speech and dialogue systems #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2304.07327
Abstract
Aligning large language models (LLMs) with human preferences has proven to drastically improve usability and has driven rapid adoption as demonstrated by ChatGPT. Alignment techniques such as supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) greatly reduce the required skill and domain knowledge to effectively harness the capabilities of LLMs, increasing their accessibility and utility across various domains. However, state-of-the-art alignment techniques like RLHF rely on high-quality human feedback data, which is expensive to create and often remains proprietary. In an effort to democratize research on large-scale alignment, we release OpenAssistant Conversations, a human-generated, human-annotated assistant-style conversation corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292 quality ratings, resulting in over 10,000 complete and fully annotated conversation trees. The corpus is a product of a worldwide crowd-sourcing effort involving over 13,500 volunteers. Models trained on OpenAssistant Conversations show consistent improvements on standard benchmarks over respective base models. We release our code and data under a fully permissive licence.
Cited by
- When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
- Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
- ShareChat: A Dataset of Chatbot Conversations in the Wild
- RecipeMasterLLM: Revisiting RoboEarth in the Era of Large Language Models
- Cost-Free Neutrality for the River Method
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Search
- D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
- Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
- Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior
- PIRA: Preference-Oriented Instruction-Tuned Reward Models with Dual Aggregation
- Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
- Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
- ISExplore:Informative Segment Selection for Efficient Personalized 3D Talking Face Generation
- AlignSurvey: A Comprehensive Benchmark for Human Preferences Alignment in Social Surveys
- KV Cache Transform Coding for Compact Storage in LLM Inference
- Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
- VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
- Comprehensive and Efficient Distillation for Lightweight Sentiment Analysis Models
- FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
- Handling Missing Responses under Cluster Dependence with Applications to Language Model Evaluation
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- LM-mixup: Text Data Augmentation via Language Model based Mixup
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning
- Planned Diffusion
- Network and Systems Performance Characterization of MCP-Enabled LLM Agents
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Towards Understanding Valuable Preference Data for Large Language Model Alignment
- MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
- Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- QoNext: Towards Next-generation QoE for Foundation Models
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
- Generating High-Quality Datasets for Code Editing via Open-Source Language Models
- Semantic Representation Attack against Aligned Large Language Models
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
- What Matters in Data for DPO?
- Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
- MedQARo: A Large-Scale Benchmark for Medical Question Answering in Romanian
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- TCIA: A Task-Centric Instruction Augmentation Method for Instruction Finetuning
- ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
- SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
- DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Cluster Topology-Driven Placement of Experts Reduces Network Traffic in MoE Inference
- Can Large Models Fool the Eye? A New Turing Test for Biological Animation
- Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
- IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
- Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
- MoKA: Mixture of Kronecker Adapters
- Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning
- On the Sustainability of AI Inferences in the Edge
- CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
- Debating AI in Archaeology: applications, implications, and ethical considerations
- LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models
- An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models
- Enhancing RLHF with Human Gaze Modeling
- Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems
- Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts
- Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
- FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
- Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
- Watermarking Autoregressive Image Generation
- Reward Models in Deep Reinforcement Learning: A Survey
- DCRM: A Heuristic to Measure Response Pair Quality in Preference Optimization
- Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
- OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
- Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory
- ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
- RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection
- SPARTA ALIGNMENT: Collectively Aligning Multiple Language Models through Combat
- From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding
- Steerable Chatbots: Exploring Personalization Control Interfaces via LLM Activation Steering
- Robust Preference Optimization via Dynamic Target Margins
- Aligning Large Language Models with Implicit Preferences from User-Generated Content
- Data Pruning by Information Maximization
- ImpRAG: Retrieval-Augmented Generation with Implicit Queries
- Taming LLMs by Scaling Learning Rates with Gradient Grouping
- Data Swarms: Optimizable Generation of Synthetic Evaluation Data
- Dataset Cartography for Large Language Model Alignment: Mapping and Diagnosing Preference Data
- LASER: Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy
- Enhancing Transformation from Natural Language to Signal Temporal Logic Using LLMs with Diverse External Knowledge
- A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs
- Improving Model Alignment Through Collective Intelligence of Open-Source LLMS
- Efficient Data Selection at Scale via Influence Distillation
- Knowledge Grafting of Large Language Models
- OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models
- Reward Model Overoptimisation in Iterated RLHF
- Scalable Valuation of Human Feedback through Provably Robust Model Alignment
- Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning
- CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models
- MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models
- Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
- Listen to the Context: Towards Faithful Large Language Models for Retrieval Augmented Generation on Climate Questions
- YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
- Think Only When You Need with Large Hybrid-Reasoning Models
- Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis
- Cross-Lingual Optimization for Language Transfer in Large Language Models
- CIE: Controlling Language Model Text Generations Using Continuous Signals
- Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks
- ProDS: Preference-oriented Data Selection for Instruction Tuning
- RAFP: Identifying LLM Lineages via Rare-Region Fingerprints
- SPIRIT: Patching Speech Language Models against Jailbreak Attacks
- Mutual-Taught for Co-adapting Policy and Reward Models
- EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMs
- WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
- LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning
- REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback
- Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
- Invisible failures in human-AI interactions
- Trust The Typical
- Synthetic Interaction Data for Scalable Personalization in Large Language Models
- Reinforcing Human Behavior Simulation via Verbal Feedback
- MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs
- AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- RAG LLMs are Not Safer: A Safety Analysis of Retrieval-Augmented Generation for Large Language Models
- MAGIC: Near-Optimal Data Attribution for Deep Learning
- From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs
- PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines
- The River Method
- MAIN: Mutual Alignment Is Necessary for instruction tuning
- PolyAlign: Conditional Human-Distribution Alignment
- Transferable text data distillation by trajectory matching
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
- REANIMATOR: Reanimate Retrieval Test Collections with Extracted and Synthetic Resources
- Dual-Difficulty Curriculum Learning for Direct Preference Optimization
- Mental model shifts in human-LLM interactions
- CARE: Multilingual Human Preference Learning for Cultural Awareness
Related