Scaling Synthetic Data Creation with 1,000,000,000 Personas
2024/06/28 by Tao Ge, Ge, Tao, Xin Chan +11 · 3 voices · 167 citations
Computer Science · Mathematics · #Business #Computer science #Data science #Human–computer interaction #Innovative Human-Technology Interaction #Mathematics #Persona #Persona Design and Applications #Scaling #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2406.20094
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2024/06/28 · openalex created_date 2024/07/02 · openalex updated_date 2026/08/05
Abstract
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce Persona Hub -- a collection of 1 billion diverse personas automatically curated from web data. These 1 billion personas (~13% of the world's total population), acting as distributed carriers of world knowledge, can tap into almost every perspective encapsulated within the LLM, thereby facilitating the creation of diverse synthetic data at scale for various scenarios. By showcasing Persona Hub's use cases in synthesizing high-quality mathematical and logical reasoning problems, instructions (i.e., user prompts), knowledge-rich texts, game NPCs and tools (functions) at scale, we demonstrate persona-driven data synthesis is versatile, scalable, flexible, and easy to use, potentially driving a paradigm shift in synthetic data creation and applications in practice, which may have a profound impact on LLM research and development.
Cited by
- Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging
- Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models
- Simulating Human Memory with Language Models
- More Is Not More: What Matters for Diversity in LLM Opinions?
- Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
- Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM
- The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
- Enhancing Diversity of LLM-Generated Educational Tasks
- TCEval: Using Thermal Comfort to Assess Cognitive and Perceptual Abilities of AI
- Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs
- LikeBench: Evaluating Subjective Likability in LLMs for Personalization
- Revisiting the Reliability of Language Models in Instruction-Following
- Do Persona-Infused LLMs Affect Performance in a Strategic Reasoning Game?
- ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems
- PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
- Understanding Down Syndrome Stereotypes in LLM-Based Personas
- Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
- SO-Bench: A Structural Output Evaluation of Multimodal LLMs
- German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
- BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
- Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
- Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
- Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
- DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
- Adapting Web Agents with Synthetic Supervision
- Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
- HaluMem: Evaluating Hallucinations in Memory Systems of Agents
- Prompting for Policy: Forecasting Macroeconomic Scenarios with Synthetic LLM Personas
- MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
- How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
- Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
- HACK: Hallucinations Along Certainty and Knowledge Axes
- LimRank: Less is More for Reasoning-Intensive Information Reranking
- A Survey on LLM Mid-Training
- AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
- QueST: Incentivizing LLMs to Generate Difficult Problems
- Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
- RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
- Merlin's Whisper: Enabling Efficient Reasoning in LLMs via Black-box Adversarial Prompting
- The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
- AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
- Aligning Large Language Models via Fully Self-Synthetic Data
- Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- Thinkquel: A Model Dedicated to Text-to-dbt Using Synthetic Data and a Span-Aware Objective
- PrimeX: A Dataset of Worldview, Opinion, and Explanation
- Spiral of Silence in Large Language Model Agents
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- Non-Collaborative User Simulators for Tool Agents
- Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
- ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
- Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
- Harnessing Synthetic Data from Generative AI for Statistical Inference
- Conf-Profile: A Confidence-Driven Reasoning Paradigm for Label-Free User Profiling
- Enhancing LLM-Based Social Bot via an Adversarial Learning Framework
- ARE: Scaling Up Agent Environments and Evaluations
- Semi-Supervised Synthetic Data Generation with Fine-Grained Relevance Control for Short Video Search Relevance Modeling
- Benchmarking and Improving LLM Robustness for Personalized Generation
- Synthesizing Attitudes, Predicting Actions (SAPA): Behavioral Theory-Guided LLMs for Ridesourcing Mode Choice Modeling
- PILOT: Steering Synthetic Data Generation with Psychological & Linguistic Output Targeting
- SCoGen: Scenario-Centric Graph-Based Synthesis of Real-World Code Problems
- Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble
- Population-Aligned Persona Generation for LLM-based Social Simulation
- PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
- Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
- NoteBar: An AI-Assisted Note-Taking System for Personal Knowledge Management
- Jointly Reinforcing Diversity and Quality in Language Model Generations
- LongCat-Flash Technical Report
- Political Ideology Shifts in Large Language Models
- Hermes 4 Technical Report
- DiscussLLM: Teaching Large Language Models When to Speak
- Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency
- IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
- MultiRef: Controllable Image Generation with Multiple Visual References
- LLMs vs. Chinese Anime Enthusiasts: A Comparative Study on Emotionally Supportive Role-Playing
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
- Large-Scale Diverse Synthesis for Mid-Training
- Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
- Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
- SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
- A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
- Imitating Mistakes in a Learning Companion AI Agent for Online Peer Learning
- GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture
Droid: A Resource Suite for AI-Generated Code Detection- Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
- ECom-Bench: Can LLM Agent Resolve Real-World E-commerce Customer Support Issues?
- DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities
- Easy Dataset: A Unified and Extensible Framework for Synthesizing LLM Fine-Tuning Data from Unstructured Documents
- Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
- Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky
- Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations
- DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain
- FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
- LeanTutor: Towards a Verified AI Mathematical Proof Tutor
- Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study
- KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model
- Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
- LLM-Based Social Simulations Require a Boundary
- MiniCPM4: Ultra-Efficient LLMs on End Devices
- Instructing Large Language Models for Low-Resource Languages: A Systematic Study for Basque
- Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor
- AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
- MotiveBench: How Far Are We From Human-Like Motivational Reasoning in Large Language Models?
- VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation
- OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
- Recycling the Web: A Method to Enhance Pre-training Data Quality and Quantity for Language Models
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- GA-S3: Comprehensive Social Network Simulation with Group Agents
- Are Economists Always More Introverted? Analyzing Consistency in Persona-Assigned LLMs
- LocalGPT: Benchmarking and Advancing Large Language Models for Local Life Services in Meituan
- Descriptive History Representations: Learning Representations by Answering Questions
- Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis
- Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution
- From Macro to Micro: Probing Dataset Diversity in Language Model Fine-Tuning
- When Harry Meets Superman: The Role of The Interlocutor in Persona-Based Dialogue Generation
- Localizing Persona Representations in LLMs
- TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis
- Bounding Box-Guided Diffusion for Synthesizing Industrial Images and Segmentation Map
- Free Lunch for User Experience: Crowdsourcing Agents for Scalable User Studies
- MEDAL: A Framework for Benchmarking LLMs as Multilingual Open-Domain Dialogue Evaluators
- Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data
- Reverse Preference Optimization for Complex Instruction Following
- Towards Conversational Development Environments: Using Theory-of-Mind and Multi-Agent Architectures for Requirements Refinement
- Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
- The Price of Format: Diversity Collapse in LLMs
- SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data
- TAG-INSTRUCT: Controlled Instruction Complexity Enhancement through Structure-based Augmentation
- NileChat: Towards Linguistically Diverse and Culturally Aware LLMs for Local Communities
- Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
- Diverse, not Short: A Length-Controlled Data Selection Strategy for Improving Response Diversity of Language Models
- ReCopilot: Reverse Engineering Copilot in Binary Analysis
- Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
- P2VA: Converting Persona Descriptions into Voice Attributes for Fair and Controllable Text-to-Speech
- ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
- SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning
- DecIF: Improving Instruction-Following through Meta-Decomposition
- IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment
- Solver-Informed RL: Grounding Large Language Models for Authentic Optimization Modeling
- Review-Instruct: A Review-Driven Multi-Turn Conversations Generation Method for Large Language Models
- A Data Synthesis Method Driven by Large Language Models for Proactive Mining of Implicit User Intentions in Tourism
- MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents
- Towards Multi-Agent Reasoning Systems for Collaborative Expertise Delegation: An Exploratory Design Study
- Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
- EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
- Evaluating Pluralism in LLMs through Latent Perspectives
- Privasis: Synthesizing the Largest "Public" Private Dataset from Scratch
- Tmax: A simple recipe for terminal agents
- Synthetic Computers at Scale for Long-Horizon Productivity Simulation
- HyPerAlign: Interpretable Personalized LLM Alignment via Hypothesis Generation
- TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
- SocialCoach: Personalized Social Skill Learning with RL-based Agentic Tutoring and Practice
- MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents
- AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
- Improving Language Model Personas via Rationalization with Psychological Scaffolds
- MatrAIx: Simulating the World with 8.3 Billion Persona Agents
- Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction
- Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale
- Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
- OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution
- APIGen-MT: Agentic Pipeline for Multi-Turn Data Generation via Simulated Agent-Human Interplay
Discussions
- AI2 used this method to help create sufficiently diverse prompts for supervised fine-tuning, which is a strong recommendation. But I kind of can't believe it works. What gives us confidence that an L [bsky, 51 points, 9 comments]
- Scaling Synthetic Data Creation with 1,000,000,000 Personas https://arxiv.org/abs/2406.20094v1 #AI #SyntheticData [bsky, 7 points, 1 comments]
- "we introduce Persona Hub a collection of 1 billion diverse personas automatically curated from web data. These 1 billion personas (~13% of the world's total population), acting as distributed carrier [bsky, 2 points, 0 comments]
Related