OpenAssistant Conversations -- Democratizing Large Language Model Alignment
2023/04/14 by Andreas Köpf, Yannic Kilcher, Köpf, Andreas +33 · 75 citations
Computer Science · #Topic Modeling #Speech and dialogue systems #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2304.07327
Abstract
Aligning large language models (LLMs) with human preferences has proven to drastically improve usability and has driven rapid adoption as demonstrated by ChatGPT. Alignment techniques such as supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF) greatly reduce the required skill and domain knowledge to effectively harness the capabilities of LLMs, increasing their accessibility and utility across various domains. However, state-of-the-art alignment techniques like RLHF rely on high-quality human feedback data, which is expensive to create and often remains proprietary. In an effort to democratize research on large-scale alignment, we release OpenAssistant Conversations, a human-generated, human-annotated assistant-style conversation corpus consisting of 161,443 messages in 35 different languages, annotated with 461,292 quality ratings, resulting in over 10,000 complete and fully annotated conversation trees. The corpus is a product of a worldwide crowd-sourcing effort involving over 13,500 volunteers. Models trained on OpenAssistant Conversations show consistent improvements on standard benchmarks over respective base models. We release our code and data under a fully permissive licence.
Cited by
- When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
- Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
- ShareChat: A Dataset of Chatbot Conversations in the Wild
- RecipeMasterLLM: Revisiting RoboEarth in the Era of Large Language Models
- Cost-Free Neutrality for the River Method
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring
- Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Search
- D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
- Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks
- Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior
- PIRA: Preference-Oriented Instruction-Tuned Reward Models with Dual Aggregation
- Go-UT-Bench: A Fine-Tuning Dataset for LLM-Based Unit Test Generation in Go
- Convergence and Stability Analysis of Self-Consuming Generative Models with Heterogeneous Human Curation
- ISExplore:Informative Segment Selection for Efficient Personalized 3D Talking Face Generation
- AlignSurvey: A Comprehensive Benchmark for Human Preferences Alignment in Social Surveys
- KV Cache Transform Coding for Compact Storage in LLM Inference
- Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs
- Comprehensive and Efficient Distillation for Lightweight Sentiment Analysis Models
- FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
- Handling Missing Responses under Cluster Dependence with Applications to Language Model Evaluation
- Ask a Strong LLM Judge when Your Reward Model is Uncertain
- LM-mixup: Text Data Augmentation via Language Model based Mixup
- ssToken: Self-modulated and Semantic-aware Token Selection for LLM Fine-tuning
- Planned Diffusion
- Network and Systems Performance Characterization of MCP-Enabled LLM Agents
- Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
- Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
- Information-Theoretic Reward Modeling for Stable RLHF: Detecting and Mitigating Reward Hacking
- Towards Understanding Valuable Preference Data for Large Language Model Alignment
- MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
- Pragyaan: Designing and Curating High-Quality Cultural Post-Training Datasets for Indian Languages
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming Vulnerability
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
- CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task
- Toward Preference-aligned Large Language Models via Residual-based Model Steering
- QoNext: Towards Next-generation QoE for Foundation Models
- TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
- Generating High-Quality Datasets for Code Editing via Open-Source Language Models
- Semantic Representation Attack against Aligned Large Language Models
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- Towards Safeguarding LLM Fine-tuning APIs against Cipher Attacks
- What Matters in Data for DPO?
- Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality
- Getting In Contract with Large Language Models -- An Agency Theory Perspective On Large Language Model Alignment
- Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding
- MedQARo: A Large-Scale Benchmark for Medical Question Answering in Romanian
- Implicit Reasoning in Large Language Models: A Comprehensive Survey
- TCIA: A Task-Centric Instruction Augmentation Method for Instruction Finetuning
- ReSURE: Regularizing Supervision Unreliability for Multi-turn Dialogue Fine-tuning
- SyGra: A Unified Graph-Based Framework for Scalable Generation, Quality Tagging, and Management of Synthetic Data
- DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
- On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- Diversity First, Quality Later: A Two-Stage Assumption for Language Model Alignment
- Cluster Topology-Driven Placement of Experts Reduces Network Traffic in MoE Inference
- Can Large Models Fool the Eye? A New Turing Test for Biological Animation
- Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
- IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
- Forgetting: A New Mechanism Towards Better Large Language Model Fine-tuning
- MoKA: Mixture of Kronecker Adapters
- Beyond the Trade-off: Self-Supervised Reinforcement Learning for Reasoning Models' Instruction Following
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- Refine-n-Judge: Curating High-Quality Preference Chains for LLM-Fine-Tuning
- On the Sustainability of AI Inferences in the Edge
- CUS-QA: Local-Knowledge-Oriented Open-Ended Question Answering Dataset
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
- Debating AI in Archaeology: applications, implications, and ethical considerations
- LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models
- An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models
- Enhancing RLHF with Human Gaze Modeling
- Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems
- Nile-Chat: Egyptian Language Models for Arabic and Latin Scripts
- Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
- Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
Related