Scaling Synthetic Data Creation with 1,000,000,000 Personas
2024/06/28 by Tao Ge, Ge, Tao, Xin Chan +11 · 3 voices · 85 citations
Computer Science · #Innovative Human-Technology Interaction #Persona Design and Applications #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2406.20094
openalex publication_date 2024/06/28 · openalex created_date 2024/07/02 · openalex updated_date 2026/07/28
Abstract
We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data. To fully exploit this methodology at scale, we introduce Persona Hub -- a collection of 1 billion diverse personas automatically curated from web data. These 1 billion personas (~13% of the world's total population), acting as distributed carriers of world knowledge, can tap into almost every perspective encapsulated within the LLM, thereby facilitating the creation of diverse synthetic data at scale for various scenarios. By showcasing Persona Hub's use cases in synthesizing high-quality mathematical and logical reasoning problems, instructions (i.e., user prompts), knowledge-rich texts, game NPCs and tools (functions) at scale, we demonstrate persona-driven data synthesis is versatile, scalable, flexible, and easy to use, potentially driving a paradigm shift in synthetic data creation and applications in practice, which may have a profound impact on LLM research and development.
Cited by
- Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging
- Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models
- Simulating Human Memory with Language Models
- More Is Not More: What Matters for Diversity in LLM Opinions?
- Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
- Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM
- The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
- Enhancing Diversity of LLM-Generated Educational Tasks
- TCEval: Using Thermal Comfort to Assess Cognitive and Perceptual Abilities of AI
- Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs
- LikeBench: Evaluating Subjective Likability in LLMs for Personalization
- Revisiting the Reliability of Language Models in Instruction-Following
- Do Persona-Infused LLMs Affect Performance in a Strategic Reasoning Game?
- ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems
- PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
- Understanding Down Syndrome Stereotypes in LLM-Based Personas
- Misalignment of LLM-Generated Personas with Human Perceptions in Low-Resource Settings
- SO-Bench: A Structural Output Evaluation of Multimodal LLMs
- German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
- BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
- Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
- Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
- Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
- DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
- Adapting Web Agents with Synthetic Supervision
- Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
- HaluMem: Evaluating Hallucinations in Memory Systems of Agents
- Prompting for Policy: Forecasting Macroeconomic Scenarios with Synthetic LLM Personas
- MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
- How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
- Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
- HACK: Hallucinations Along Certainty and Knowledge Axes
- LimRank: Less is More for Reasoning-Intensive Information Reranking
- A Survey on LLM Mid-Training
- AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
- QueST: Incentivizing LLMs to Generate Difficult Problems
- Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
- RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
- Merlin's Whisper: Enabling Efficient Reasoning in LLMs via Black-box Adversarial Prompting
- The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation
- AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratch
- Aligning Large Language Models via Fully Self-Synthetic Data
- Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
- Thinkquel: A Model Dedicated to Text-to-dbt Using Synthetic Data and a Span-Aware Objective
- PrimeX: A Dataset of Worldview, Opinion, and Explanation
- Spiral of Silence in Large Language Model Agents
- SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents
- Non-Collaborative User Simulators for Tool Agents
- Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data
- ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning
- Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
- BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
- HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
- Harnessing Synthetic Data from Generative AI for Statistical Inference
- Conf-Profile: A Confidence-Driven Reasoning Paradigm for Label-Free User Profiling
- Enhancing LLM-Based Social Bot via an Adversarial Learning Framework
- ARE: Scaling Up Agent Environments and Evaluations
- Semi-Supervised Synthetic Data Generation with Fine-Grained Relevance Control for Short Video Search Relevance Modeling
- Benchmarking and Improving LLM Robustness for Personalized Generation
- Synthesizing Attitudes, Predicting Actions (SAPA): Behavioral Theory-Guided LLMs for Ridesourcing Mode Choice Modeling
- PILOT: Steering Synthetic Data Generation with Psychological & Linguistic Output Targeting
- SCoGen: Scenario-Centric Graph-Based Synthesis of Real-World Code Problems
- Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble
- Population-Aligned Persona Generation for LLM-based Social Simulation
- PersonaFuse: A Personality Activation-Driven Framework for Enhancing Human-LLM Interactions
- Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
- NoteBar: An AI-Assisted Note-Taking System for Personal Knowledge Management
- Jointly Reinforcing Diversity and Quality in Language Model Generations
- LongCat-Flash Technical Report
- Political Ideology Shifts in Large Language Models
- Hermes 4 Technical Report
- DiscussLLM: Teaching Large Language Models When to Speak
- Input-Time Scaling
- IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
- MultiRef: Controllable Image Generation with Multiple Visual References
- LLMs vs. Chinese Anime Enthusiasts: A Comparative Study on Emotionally Supportive Role-Playing
- CUPID: Evaluating Personalized and Contextualized Alignment of LLMs from Interactions
- The Art of Breaking Words: Rethinking Multilingual Tokenizer Design
- Large-Scale Diverse Synthesis for Mid-Training
- Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Training
- Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing Styles
- SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment
Discussions
- AI2 used this method to help create sufficiently diverse prompts for supervised fine-tuning, which is a strong recommendation. But I kind of can't believe it works. What gives us confidence that an L [bsky, 51 points, 9 comments]
- Scaling Synthetic Data Creation with 1,000,000,000 Personas https://arxiv.org/abs/2406.20094v1 #AI #SyntheticData [bsky, 7 points, 1 comments]
- "we introduce Persona Hub a collection of 1 billion diverse personas automatically curated from web data. These 1 billion personas (~13% of the world's total population), acting as distributed carrier [bsky, 2 points, 0 comments]
Related