LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
2023/09/21 by Lianmin Zheng, Wei-Lin Chiang, Zheng, Lianmin +24 · 6 voices · 155 citations
Computer Science · Psychology · #Benchmark (surveying) #Cartography #Computer science #Conversation #Data science #Geography #Natural Language Processing Techniques #Process (computing) #Psychology #Resource (disambiguation) #Scale (ratio) #Text Readability and Simplification #Topic Modeling #World Wide Web
paper · pdf · doi:10.48550/arxiv.2309.11998
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/09/21 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Studying how people interact with large language models (LLMs) in real-world scenarios is increasingly important due to their widespread use in various applications. In this paper, we introduce LMSYS-Chat-1M, a large-scale dataset containing one million real-world conversations with 25 state-of-the-art LLMs. This dataset is collected from 210K unique IP addresses in the wild on our Vicuna demo and Chatbot Arena website. We offer an overview of the dataset's content, including its curation process, basic statistics, and topic distribution, highlighting its diversity, originality, and scale. We demonstrate its versatility through four use cases: developing content moderation models that perform similarly to GPT-4, building a safety benchmark, training instruction-following models that perform similarly to Vicuna, and creating challenging benchmark questions. We believe that this dataset will serve as a valuable resource for understanding and advancing LLM capabilities. The dataset is publicly available at https://huggingface.co/datasets/lmsys/lmsys-chat-1m.
Cited by
- What AI Red-Team Evaluations Can and Cannot Prove
- After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions
- A CXL Memory Rack for Multi-Turn LLM Serving
- Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
- KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
- SODA: Semi On-Policy Black-Box Distillation for Large Language Models
- Align AI to Dynamic Human-AI Workflows
- AI Fiction in the Wild
- Economic Evaluations of Language Models
- Greater accessibility can amplify discrimination in generative AI
- Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
- LLMs Get Lost In Multi-Turn Conversation
- Sycophantic AI decreases prosocial intentions and promotes dependence
- BitNet b1.58 2B4T Technical Report
- Idiosyncrasies in Large Language Models
- Why human-AI relationships need socioaffective alignment
- Splitwise: Collaborative Edge-Cloud Inference for LLMs via Lyapunov-Assisted DRL
- The One-Word Census: Answer-Choice Conformity Across 44 Language Models
- Learning from 53.6K Real-World Developer Edits of AI-Generated Code
- Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
- ShareChat: A Dataset of Chatbot Conversations in the Wild
- Affect, Body, Cognition, Demographics, and Emotion: The ABCDE of Text Features for Computational Affective Science
- Activation Oracles: Training and Evaluating LLMs as General-Purpose Activation Explainers
- Towards Effective Model Editing for LLM Personalization
- Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
- Decoding Human-LLM Collaboration in Coding: An Empirical Study of Multi-Turn Conversations in the Wild
- Log Probability Tracking of LLM APIs
- Unsupervised decoding of encoded reasoning using language model interpretability
- The PLLuM Instruction Corpus
- HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA
- Generative Caching for Structurally Similar Prompts and Responses
- Black-Box On-Policy Distillation of Large Language Models
- Sim4IA-Bench: A User Simulation Benchmark Suite for Next Query and Utterance Prediction
- Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
- EASE: Practical and Efficient Safety Alignment for Small Language Models
- Unsupervised Evaluation of Multi-Turn Objective-Driven Interactions
- Verifying LLM Inference to Prevent Model Weight Exfiltration
- ARC-GEN: A Mimetic Procedural Benchmark Generator for the Abstraction and Reasoning Corpus
- Finding Diamonds in Conversation Haystacks: A Benchmark for Conversational Data Retrieval
- Larger and more instructable language models become less reliable
- The Art of Asking: Multilingual Prompt Optimization for Synthetic Data
- SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
- Network and Systems Performance Characterization of MCP-Enabled LLM Agents
- A Survey on Evaluation of Large Language Models
- What Generative Search Engines Like and How to Optimize Web Content Cooperatively
- Verifying Chain-of-Thought Reasoning via Its Computational Graph
- Artificial Impressions: Evaluating Large Language Model Behavior Through the Lens of Trait Impressions
- LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
- The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs
- Investigating Thematic Patterns and User Preferences in LLM Interactions using BERTopic
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
- Flipping the Dialogue: Training and Evaluating User Language Models
- EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
- Auditing Pay-Per-Token in Large Language Models
- A global log for medical AI
- Atomic Thinking of LLMs: Decoupling and Exploring Mathematical Reasoning Abilities
- Non-Collaborative User Simulators for Tool Agents
- DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning
- MMPB: It's Time for Multi-Modal Personalization
- MIRAGE: Multi-hop Reasoning with Ambiguity Evaluation for Illusory Questions
- Prompt-Aware Scheduling for Low-Latency LLM Serving
- GORGO: Online Tuning for Cross-Region Network-Aware LLM Serving
- GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing
- Towards mitigating information leakage when evaluating safety monitors
- Understanding Prompt Management in GitHub Repositories: A Call for Best Practices
- Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality
- Established Psychometric vs. Ecologically Valid Questionnaires: Rethinking Psychological Assessments in Large Language Models
- IPR: Intelligent Prompt Routing with User-Controlled Quality-Cost Trade-offs
- VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing
- Improving Aviation Safety Analysis: Automated HFACS Classification Using Reinforcement Learning with Group Relative Policy Optimization
- Generative Interfaces for Language Models
- Adaptively Robust LLM Inference Optimization under Prediction Uncertainty
- Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
- Intent-Aware Schema Generation And Refinement For Literature Review Tables
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
- Conversational Inoculation to Enhance Resistance to Misinformation
- LLM Serving Optimization with Variable Prefill and Decode Lengths
- LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
- Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
- MedSynth: Realistic, Synthetic Medical Dialogue-Note Pairs
- Show or Tell? Modeling the evolution of request-making in Human-LLM conversations
- Exploring Direct Instruction and Summary-Mediated Prompting in LLM-Assisted Code Modification
- Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
- TweakLLM: A Routing Architecture for Dynamic Tailoring of Cached Responses
- LeMix: Unified Scheduling for LLM Training and Inference on Multi-GPU Systems
- PurpCode: Reasoning for Safer Code Generation
- Checklists Are Better Than Reward Models For Aligning Language Models
- Cybersecurity AI (CAI) Dataset
- PolyServe: Efficient Multi-SLO Serving at Scale
- Learning to summarize user information for personalized reinforcement learning from human feedback
- Toward Real-World Chinese Psychological Support Dialogues: CPsDD Dataset and a Co-Evolving Multi-Agent System
- Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report
- DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities
- Mass-Scale Analysis of In-the-Wild Conversations Reveals Complexity Bounds on LLM Jailbreaking
- Towards Resource-Efficient Serverless LLM Inference with SLINFER
- TaP: A Taxonomy-Guided Framework for Automated and Scalable Preference Data Generation
- Agent.xpu: Efficient Scheduling of Agentic LLM Workloads on Heterogeneous SoC
- Persona Features Control Emergent Misalignment
- LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
- PARALLELPROMPT: Extracting Parallelism from Large Language Model Queries
- Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
- Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models
- ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
- Arch-Router: Aligning LLM Routing with Human Preferences
- PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech Interaction
- What Makes a Good Natural Language Prompt?
- SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs
- Search Arena: Analyzing Search-Augmented LLMs
- Cascadia: An Efficient Cascade Serving System for Large Language Models
- SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models
- From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding
- REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM
- ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratch
- Quantitative LLM Judges
- Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models
- From Chat Logs to Collective Insights: Aggregative Question Answering
- Evaluating the Sensitivity of LLMs to Prior Context
- SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference
- Principled Content Selection to Generate Diverse and Personalized Multi-Document Summaries
- Enhancing Transformation from Natural Language to Signal Temporal Logic Using LLMs with Diverse External Knowledge
- Is Your LLM Overcharging You? Tokenization, Transparency, and Incentives
- Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
- Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data Scheduling
- Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries
- DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
- Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
- MPO: Multilingual Safety Alignment via Reward Gap Optimization
- HBO: Hierarchical Balancing Optimization for Fine-Tuning Large Language Models
- ExpertSteer: Intervening in LLMs through Expert Knowledge
- RH-RAG: Trustworthy Long-Form Generation for Privacy-Constrained Settings
- EdgeWisePersona: A Dataset for On-Device User Profiling from Natural Language Interactions
- HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages
- Review-Instruct: A Review-Driven Multi-Turn Conversations Generation Method for Large Language Models
- ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Can LLMs Reliably Self-Report Adversarial Prefills, and How?
- OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations
- Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM
- Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items
- Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation
- SAGE: A Generic Framework for LLM Safety Evaluation
- ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions
- Flow-Controlled Scheduling for LLM Inference with Provable Stability Guarantees
- Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests
- JITServe: SLO-aware LLM Serving with Imprecise Request Information
- MatrAIx: Simulating the World with 8.3 Billion Persona Agents
- PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
- Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
- PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines
- Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
- DICE: A Framework for Dimensional and Contextual Evaluation of Language Models
- Mitigating Many-Shot Jailbreaking
- Efficient LLM Serving on Hybrid Real-time and Best-effort Requests
- AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-time Computation
- Societal Impacts Research Requires Benchmarks for Creative Composition Tasks
Discussions
- What do people actually talk about with LLMs? Here are the main topics, according to a new paper: arxiv.org/abs/2309.11998. A lot of questions about programming. Erotic storytelling is only #9.
Mostl [bsky, 20 points, 3 comments]
- Nearly 10% of people ask AI chatbots for explicit content. Will it lead LLMs astray? [Article from October 3] [lemmy, 15 points, 17 comments]
- Lmsys-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset [hn, 2 points, 1 comments]
- ah i mean all the uses for stuff that's not (academic) research. like Fig 3 here (2024?). there's.. lots, i guess. arxiv.org/pdf/2309.11998 [bsky, 1 points, 1 comments]
- 10 percent of all conversations with an AI/LLMs are ... erotic in nature. Know we know that. arxiv.org/abs/2309.11998 [bsky, 0 points, 0 comments]
- 📝 LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset 📚👾 "LMSYS-Chat-1M is a dataset containing one million human-chatbot conversations with 25 state-of-the-art language models, collec [mastodon, 0 points, 0 comments]
Related