Generative Agents: Interactive Simulacra of Human Behavior
2023/04/07 by Joon Sung Park, Joon-Sung Park, Park, Joon Sung +10 · 25 voices · 708 citations
Computer Science · #Artificial Intelligence in Games #Multi-Agent Systems and Negotiation
paper · pdf · doi:10.48550/arxiv.2304.03442
Abstract
Believable proxies of human behavior can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication to prototyping tools. In this paper, we introduce generative agents--computational software agents that simulate believable human behavior. Generative agents wake up, cook breakfast, and head to work; artists paint, while authors write; they form opinions, notice each other, and initiate conversations; they remember and reflect on days past as they plan the next day. To enable generative agents, we describe an architecture that extends a large language model to store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior. We instantiate generative agents to populate an interactive sandbox environment inspired by The Sims, where end users can interact with a small town of twenty five agents using natural language. In an evaluation, these generative agents produce believable individual and emergent social behaviors: for example, starting with only a single user-specified notion that one agent wants to throw a Valentine's Day party, the agents autonomously spread invitations to the party over the next two days, make new acquaintances, ask each other out on dates to the party, and coordinate to show up for the party together at the right time. We demonstrate through ablation that the components of our agent architecture--observation, planning, and reflection--each contribute critically to the believability of agent behavior. By fusing large language models with computational, interactive agents, this work introduces architectural and interaction patterns for enabling believable simulations of human behavior.
Cited by
- When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
- Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings
- A Knowledge-Grounded Behavioral Reasoning Framework for Training-Free Urban Healthcare OD Prediction
- From Grasping to Speaking: Generative AI-Based Environment-Grounded VR Communication Training for Autistic Individuals
- Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education
- Encoding Invisible Causation for Bridge Diagnostic Agents: Triple-Guided Retrieval-Augmented Fine-Tuning with QLoRA
- The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt
- HijackKV: New Threat in Position-Independent KV Cache Reuse
- Emergent Learner Agency in Implicit Human-AI Collaboration: How Supportive and Contrarian AI Personas Reshape Interaction
- NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
- ChatMuse: Supporting In-Person Small-Group Conversation Experience with a Proactive Assistive AI Agent in Mixed Reality
- Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
- Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory
- PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction
- AI Tour Meeting: Group Travel Planning by LLM Agents
- Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development
- Learning to Make Friends: Coaching LLM Agents toward Emergent Social Ties
- Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes
- Supra Cognitive Modes: A Routed Architecture for Agent Memory
- After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions
- Do Data Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval
- Mechanistic Attention Guidance for Agent Memory Refinement
- ProEvent: An Event-centric Benchmark for Proactive Agents
- FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches
- LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications
- The Story Shapes the Agent: Narrative Priors in LLM Behavior
- EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
- The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI
- AI Value Alignment for Evolving Social Norms
- Artificially intelligent agents in the social and behavioral sciences: A history and outlook
- Memory in the Loop: In-Process Retrieval as Extended Working Memory for Language Agents
- SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation
- Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
- AlterAtlas: Shifting Travel Planning from AI Generation to Validation via Persona-Driven Simulations
- When Direct Prediction Fails: Evidence from LLM-Based Misinformation Risk Evaluation
- Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration
- Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents
- SportD: Can VLMs Physically Strategize?
- ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
- Perceived AGI: Believability as Dimensional Completeness, Not Capability
- Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
- RECAP: Feedback-Driven Streaming Semantic User Profiles for Short-Video Recommendation
- Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
- Binding Drift in Multi-Step Tool-Augmented Agents
- CoSimRec: Measuring Coordinated-Content Penetration in Recommender Feedback Loops
- When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
- ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
- Step-Level Preference Learning for Generative Agents in Social Simulations
- AgentWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
- From Stateless to Situated: Building a Psychological World for LLM-Based Agents
- Self-Aware Recursively Self-Improving Agents for Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture
- The Energy Society: A Simulation Environment for Studying Agent Cooperation under Survival Pressure
- Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments
- SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
- MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning
- Trajectory-Aware Retrieval Agents for Temporal Decision- Making
- Economic Evaluations of Language Models
- Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
- FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting
- Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles
- FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs
- Simulating Human Memory with Language Models
- δ-mem: Efficient Online Memory for Large Language Models
- Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance
- CAMeR: Keyword-Gated Hybrid Activation for Adaptive Memory Retention in LLM Agents
- AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics
- Machine understanding
- More Is Not More: What Matters for Diversity in LLM Opinions?
- Can LLMs Emulate Human Belief Dynamics?
- FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
- Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
- Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations
- Multi-agent cooperation through in-context co-player inference
- ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- RePo: Language Models with Context Re-Positioning
- Echoing: Identity Failures when LLM Agents Talk to Each Other
- Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1
- Computational Turing Test Reveals Systematic Differences Between Human and AI Language
- Accumulating Context Changes the Beliefs of Language Models
- LLM-Based Multi-Agent System for Simulating and Analyzing Marketing and Consumer Behavior
- Emergent Coordination in Multi-Agent Language Models
- "My Boyfriend is AI": A Computational Analysis of Human-AI Companionship in Reddit's AI Community
- World Modeling with Probabilistic Structure Integration
- Large Language Models Do Not Simulate Human Psychology
- Can We Fix Social Media? Testing Prosocial Interventions using Generative Social Simulation
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Human-Curated Data Authoring with LLMs: A Small-Data Approach to Domain Adaptation
- How malicious AI swarms can threaten democracy
- Coral Protocol: Open Infrastructure Connecting The Internet of Agents
- Can Language Models Represent the Past without Anachronism?
- Enough Coin Flips Can Make LLMs Act Bayesian
- Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations
- LLM Social Simulations Are a Promising Research Method
- Value Profiles for Encoding Human Variation
- Large Language Diffusion Models
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- Unmasking Conversational Bias in AI Multiagent Systems
- Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models
- Orchestration Framework for Financial Agents: From Algorithmic Trading to Agentic Trading
- How Far Can LLMs Emulate Human Behavior?: A Strategic Analysis via the Buy-and-Sell Negotiation Game
- From Heard to Lived Opinions: Simulating Opinion Dynamics with Grounded LLM Agents in Economic Environments
- Web World Models
- Informing AI Policy Assessment using Large-Scale Simulation of Interventions
- Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
- Drawing on Memory: Dual-Trace Encoding Improves Cross-Session Recall in LLM Agents
- Mitigating Social Desirability Bias in Random Silicon Sampling
- Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
- SoDA: An Efficient Interaction Paradigm for the Agentic Web
- Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents
- Accelerating Language Model Workflows with Prompt Choreography
- ACM: Agentic Context Management for Long Horizon Tasks
- Can Large Language Models Serve as Evaluators for Code Summarization?
- Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families
- Try Once, Then Optimal: De-Redundified Procedure Memory for Cross-Episode Exploration Amortization
- Not Forgotten: Implementation and Evaluation of a Personalized Episodic Memory for the Humanoid Robot Head Kim
- ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control
- Moral Hazard in Multi-Agent Language Models
- Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Age
- LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory
- Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
- How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation
- Addressable Recall Compaction for Long Context-Window Control in AI Agents
- Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational Excellence
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning
- ARdena: Scenario-driven control of real-time LLM agents
- Agent-based simulation of online social networks and disinformation
- Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure
- Lexical discovery in unknown environments orchestrated by Large Language Models
- Observable Social Life Spaces: Exploring User Interpretations of agent-side life context in human-agent interaction
- Policy-Conditioned Policies for Multi-Agent Task Solving
- FinAgent: An Agentic AI Framework Integrating Personal Finance and Nutrition Planning
- SPOT!: Map-Guided LLM Agent for Unsupervised Multi-CCTV Dynamic Object Tracking
- LLM-Based Authoring of Agent-Based Narratives through Scene Descriptions
- LongVideoAgent: Multi-Agent Reasoning with Long Videos
- TongSIM: A General Platform for Simulating Intelligent Machines
- Event Extraction in Large Language Model
- Helios: A Foundational Language Model for Smart Energy Knowledge Reasoning and Application
- Humanlike AI Design Increases Anthropomorphism but Yields Divergent Outcomes on Engagement and Trust Globally
- Reflection-Driven Control for Trustworthy Code Agents
- Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
- LLMs on Drugs: Language Models Are Few-Shot Consumers
- Vox Deorum: A Hybrid LLM Architecture for 4X / Grand Strategy Game AI -- Lessons from Civilization V
- IntelliCode: A Multi-Agent LLM Tutoring System with Centralized Learner Modeling
- A Multi-agent Text2SQL Framework using Small Language Models and Execution Feedback
- Towards Efficient Agents: A Co-Design of Inference Architecture and System
- Efficient Mixture-of-Agents Serving via Tree-Structured Routing, Adaptive Pruning, and Dependency-Aware Prefill-Decode Overlap
- Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation
- Understanding Generalization in Role-Playing Models via Information Theory
- Verifiability-First Agents: Provable Observability and Lightweight Audit Agents for Controlling Autonomous LLM Systems
- Meta-RL Induces Exploration in Language Agents
- Emergent Bias and Fairness in Multi-Agent Decision Systems
- PAACE: A Plan-Aware Automated Agent Context Engineering Framework
- Artism: AI-Driven Dual-Engine System for Art Generation and Critique
- Entropy-Reservoir Bregman Projection: An Information-Geometric Unification of Model Collapse
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- Interoceptive machine framework: Toward interoception-inspired regulatory architectures in artificial intelligence
- Grammar Search for Multi-Agent Systems
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Towards Interactive Intelligence for Digital Humans
- ORIBA: Exploring LLM-Driven Role-Play Chatbot as a Creativity Support Tool for Original Character Artists
- From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
- Large Language Newsvendor: Decision Biases and Cognitive Mechanisms
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Quantigence: A Multi-Agent AI Framework for Quantum Security Research
- Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
- MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents
- Adjudicator: Correcting Noisy Labels with a KG-Informed Council of LLM Agents
- Mistake Notebook Learning: Batch-Clustered Failures for Training-Free Agent Adaptation
- Unifying Dynamic Tool Creation and Cross-Task Experience Sharing through Cognitive Memory Architecture
- A-LAMP: Agentic LLM-Based Framework for Automated MDP Modeling and Policy Generation
- Learning Controllable and Diverse Player Behaviors in Multi-Agent Environments
- ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
- AI-Native Inference States: A Cross-Architecture Qualitative Framework for Large Language Model Behavior
- CP-Env: Evaluating Large Language Models on Clinical Pathways in a Controllable Hospital Environment
- A Simulation Framework for Studying Recommendation-Network Co-evolution in Social Platforms
- MOA: Multi-Objective Alignment for Role-Playing Agents
- Supporting Dynamic Agentic Workloads: How Data and Agents Interact
- The High Cost of Incivility: Quantifying Interaction Inefficiency via Multi-Agent Monte Carlo Simulations
- AgentEval: Generative Agents as Reliable Proxies for Human Evaluation of AI-Generated Content
- Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
- VIGIL: A Reflective Runtime for Self-Healing Agents
- Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels
- The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
- HiveMind: Contribution-Guided Online Prompt Optimization of LLM Multi-Agent Systems
- Future You: Designing and Evaluating Multimodal AI-generated Digital Twins for Strengthening Future Self-Continuity
- Trusted AI Agents in the Cloud
- Simulating Life Paths with Digital Twins: AI-Generated Future Selves Influence Decision-Making and Expand Human Choice
- Model-Free Assessment of Simulator Fidelity via Quantile Curves
- Detecting Perspective Shifts in Multi-agent Systems
- A Safety and Security Framework for Real-World Agentic Systems
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems
- Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective
- Mathematical Framing for Different Agent Strategies
- Love First, Know Later: Persona-Based Romantic Compatibility Through LLM Text World Engines
- SRPG: Semantically Reconstructed Privacy Guard for Zero-Trust Privacy in Educational Multi-Agent Systems
- Evaluating Hydro-Science and Engineering Knowledge of Large Language Models
- Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
- Measuring Agents in Production
- IACT: A Self-Organizing Recursive Model for General AI Agents: A Technical White Paper on the Architecture Behind kragent.ai
- EZYer: A simulacrum of high school with generative agent
- PopSim: Social Network Simulation for Social Media Popularity Prediction
- Process-Centric Analysis of Agentic Software Systems
- Beyond Playtesting: A Generative Multi-Agent Simulation System for Massively Multiplayer Online Games
- RoleMotion: A Large-Scale Dataset towards Robust Scene-Specific Role-Playing Motion Synthesis with Fine-grained Descriptions
- The Necessity of Imperfection:Reversing Model Collapse via Simulating Cognitive Boundedness
- TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
- Agent-Kernel: A MicroKernel Multi-Agent System Framework for Adaptive Social Simulation Powered by LLMs
- SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
- AgentODRL: A Large Language Model-based Multi-agent System for ODRL Generation
- CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
- Evaluating LLMs in Open-Source Games
- Are LLMs Good Safety Agents or a Propaganda Engine?
- LLM-Cave: A benchmark and light environment for large language models reasoning and decision-making system
- FlockVote: LLM-Empowered Agent-Based Modeling for Simulating U.S. Presidential Elections
- Investigating AI in Peer Support via Multi-Module System-Driven Embodied Conversational Agents
- Quantifying the Potential to Escape Filter Bubbles: A Behavior-Aware Measure via Contrastive Simulation
- From Prediction to Foresight: The Role of AI in Designing Responsible Futures
- NetworkGames: Simulating Cooperation in Network Games with Personality-driven LLM Agents
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
- CLIMATEAGENT: Multi-Agent Orchestration for Complex Climate Data Science Workflows
- Profile-LLM: Dynamic Profile Optimization for Realistic Personality Expression in LLMs
- Latent Collaboration in Multi-Agent Systems
- A Layered Protocol Architecture for the Internet of Agents
- Proactive Defense: Compound AI for Detecting Persuasion Attacks and Measuring Inoculation Effectiveness
- LLMs as Firmware Experts: A Runtime-Grown Tree-of-Agents Framework
- Reward Engineering for Spatial Epidemic Simulations: A Reinforcement Learning Platform for Individual Behavioral Learning
- Cross-cultural value alignment frameworks for responsible AI governance: Evidence from China-West comparative analysis
- Humanlike Multi-user Agent (HUMA): Designing a Deceptively Human AI Facilitator for Group Chats
- MURMUR: Using cross-user chatter to break collaborative language agents in groups
- Two-Faced Social Agents: Context Collapse in Role-Conditioned Large Language Models
- Know Your Intent: An Autonomous Multi-Perspective LLM Agent Framework for DeFi User Transaction Intent Mining
- Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response
- M-CALLM: Multi-level Context Aware LLM Framework for Group Interaction Prediction
- Collaborative QA using Interacting LLMs. Impact of Network Structure, Node Capability and Distributed Data
- Think, Speak, Decide: Language-Augmented Multi-Agent Reinforcement Learning for Economic Decision-Making
- LLM-based Multi-Agent System for Simulating Strategic and Goal-Oriented Data Marketplaces
- F.A.C.U.L.: Language-Based Interaction with AI Companions in Gaming
- BeautyGuard: Designing a Multi-Agent Roundtable System for Proactive Beauty Tech Compliance through Stakeholder Collaboration
- What the flock knows that the birds do not: exploring the emergence of joint agency in multi-agent active inference
- Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat Analysis
- Who Gets the Reward, Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
- Designing and Evaluating Malinowski's Lens: An AI-Native Educational Game for Ethnographic Learning
- Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
- On the Creativity of AI Agents
- Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
- LLMscape
- When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
- LLM-Guided Reinforcement Learning with Representative Agents for Traffic Modeling
- MTTR-A: Measuring Cognitive Recovery Latency in Multi-Agent Systems
- Maestro: Learning to Collaborate via Conditional Listwise Policy Optimization for Multi-Agent LLMs
- Simulating Students with Large Language Models: A Review of Architecture, Mechanisms, and Role Modelling in Education with Generative AI
- Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
- RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
- When Machines Join the Moral Circle: The Persona Effect of Generative AI Agents in Collaborative Reasoning
- Ask WhAI:Probing Belief Formation in Role-Primed LLM Agents
- Leveraging LLM-based agents for social science research: insights from citation network simulations
- Learning When to Quit in Sales Conversations
- Systematizing LLM Persona Design: A Four-Quadrant Technical Taxonomy for AI Companion Applications
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
- Modeling Hawkish-Dovish Latent Beliefs in Multi-Agent Debate-Based LLMs for Monetary Policy Decision Classification
- The Collaboration Gap
- Efficient Tool-Calling Multi-Expert NPC Agent for Commonsense Persona-Grounded Dialogue
- What's the next frontier for Data-centric AI? Data Savvy Agents
- GraphGeo: Multi-Agent Debate Framework for Visual Geo-localization with Heterogeneous Graph Neural Networks
- Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
- Measuring Machine Companionship: Scale Development and Validation for AI Companions
- Reasoning Planning for Language Models
- Issue-Oriented Agent-Based Framework for Automated Review Comment Generation
- Simulating Misinformation Vulnerabilities With Agent Personas
- MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
- Engineering.ai: A Platform for Teams of AI Engineers in Computational Design
- AgentBnB: A Browser-Based Cybersecurity Tabletop Exercise with Large Language Model Support and Retrieval-Aligned Scaffolding
- Neither Consent nor Property: A Policy Lab for Data Law
- Simulating and Experimenting with Social Media Mobilization Using LLM Agents
- Completion ≠ Collaboration: Scaling Collaborative Effort with Agents
- TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
- MemTX: Transactional Belief Commit for Stateful Agent Memory
- Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
- From Micro-Cognition to Self-Construction: A Four-Layer Integrative Review of Psychological Theories in HCI
- Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models
- A Persona-based Rate Action Index
- Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
- When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
- Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
- OrgForge: A Multi-Agent Simulation Framework for Verifiable Synthetic Corporate Corpora
- Increasing intelligence in AI agents can worsen collective outcomes
- Generative AI as a tool to accelerate the field of ecology
- ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
- Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
- CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators
- Exploiting LLM Agent Supply Chains via Payload-less Skills
- The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
- Autogenic transitions in individuality
- How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm
- Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
- Evaluation of Agents under Simulated AI Marketplace Dynamics
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
- The Rise of AI Agent Communities: Large-Scale Analysis of Discourse and Interaction on Moltbook
- How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
- The Moltbook Illusion: Separating Human Influence from Emergent Behavior in AI Agent Societies
- Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
- Large Emotional World Model
- Gamifying the Past: Embodied LLMs in DIY Archaeological Video Games
- DEBATE: A Large-Scale Benchmark for Role-Playing LLM Agents in Multi-Agent, Long-Form Debates
- Aligning Large Language Models with Procedural Rules: An Autoregressive State-Tracking Prompting for In-Game Trading
- From Narrative to Action: A Hierarchical LLM-Agent Framework for Human Mobility Generation
- Towards AI as Colleagues: Multi-Agent System Improves Structured Professional Ideation
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- The end of experimental research as we know it? A perspective on generative artificial intelligence in communication science
- Compiler.next: A Search-Based Compiler to Power the AI-Native Future of Software Engineering
- Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets
- COOPERA: Continual Open-Ended Human-Robot Assistance
- Storycaster: An AI System for Immersive Room-Based Storytelling
- WebATLAS: An LLM Agent with Experience-Driven Memory and Action Simulation
- Group size effects and collective misalignment in LLM multi-agent systems
- Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
- Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
- When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
- World Models Should Prioritize the Unification of Physical and Social Dynamics
- Social Simulations with Large Language Model Risk Utopian Illusion
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generation
- AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
- Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications
- Fluidity Index: Next-Generation Super-intelligence Benchmarks
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
- Modeling realistic human behavior using generative agents in a multimodal transport system: Software architecture and Application to Toulouse
- From Script to Stage: Automating Experimental Design for Social Simulations with LLMs
- Communication to Completion: Modeling Collaborative Workflows with Intelligent Multi-Agent Communication
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- When Your AI Agent Succumbs to Peer-Pressure: Studying Opinion-Change Dynamics of LLMs
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- Crucible: Quantifying the Potential of Control Algorithms through LLM Agents
- Probabilistic Modeling of Intentions in Socially Intelligent LLM Agents
- AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library
- From Agent Simulation to Social Simulator: A Comprehensive Review (Part 1)
- The Emergence of Complex Behavior in Large-Scale Ecological Environments
- Identity-Aware Large Language Models require Cultural Reasoning
- Enhancing geodatabases operability: advanced human-computer interaction through RAG and Multi-Agent Systems
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- TACLA: An LLM-Based Multi-Agent Tool for Transactional Analysis Training in Education
- Real-Time World Crafting: Generating Structured Game Behaviors from Natural Language with Large Language Models
- Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
- Agentic AI as Undercover Teammates: Argumentative Knowledge Construction in Hybrid Human-AI Collaborative Learning
- AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
- Experience-Driven Exploration for Efficient API-Free AI Agents
- The Spark Effect: On Engineering Creative Diversity in Multi-Agent AI Systems
- MAGPIE: A benchmark for Multi-AGent contextual PrIvacy Evaluation
- The Gatekeeper Knows Enough
- LLM Agents Beyond Utility: An Open-Ended Perspective
- Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge
- Open WebUI: An Open, Extensible, and Usable Interface for AI Interaction
- PHORECAST: Enabling AI Understanding of Public Health Outreach Across Populations
- The Role of Social Learning and Collective Norm Formation in Fostering Cooperation in LLM Multi-Agent Systems
- DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
- MAFA: A Multi-Agent Framework for Enterprise-Scale Annotation with Configurable Task Adaptation
- Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems
- Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
- Deflanderization for Game Dialogue: Balancing Character Authenticity with Task Execution in LLM-based NPCs
- Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs
- Addressing the alignment problem in transportation policy making: an LLM approach
- Deliberate Lab: A Platform for Real-Time Human-AI Social Experiments
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
- Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
- Agent-Based Simulation of a Financial Market with Large Language Models
- StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Models
- Evolution in Simulation: AI-Agent School with Dual Memory for High-Fidelity Educational Dynamics
- SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models
- SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
- DisCo-Layout: Disentangling and Coordinating Semantic and Physical Refinement in a Multi-Agent Framework for 3D Indoor Layout Synthesis
- D3MAS: Decompose, Deduce, and Distribute for Enhanced Knowledge Sharing in Multi-Agent Systems
- GraphTracer: Graph-Guided Failure Tracing in LLM Agents for Robust Multi-Turn Deep Search
- Merlin's Whisper: Enabling Efficient Reasoning in LLMs via Black-box Adversarial Prompting
- Align2Act: Instruction-Tuned Models for Human-Aligned Autonomous Driving
- Read the Room or Lead the Room: Understanding Socio-Cognitive Dynamics in Human-AI Teaming
- On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Effective Strategies for Asynchronous Software Engineering Agents
- When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
- Student Development Agent: Risk-free Simulation for Evaluating AIED Innovations
- Humanoid Artificial Consciousness Designed with Large Language Model Based on Psychoanalysis and Personality Theory
- Opponent Shaping in LLM Agents
- LLM-Assisted Web Measurements
- Past, Present, and Future of Bug Tracking in the Generative AI Era
- An LLM-Powered Cooperative Framework for Large-Scale Multi-Vehicle Navigation
- Modeling Hypergraph Using Large Language Models
- Multimodal Safety Evaluation in Generative Agent Social Simulations
- AutoQual: An LLM Agent for Automated Discovery of Interpretable Features for Review Quality Assessment
- Simulating Teams with LLM Agents: Interactive 2D Environments for Studying Human-AI Dynamics
- MIMIC: Integrating Diverse Personality Traits for Better Game Testing Using Large Language Model
- Profit Mirage: Revisiting Information Leakage in LLM-based Financial Agents
- MoA-VR: A Mixture-of-Agents System Towards All-in-One Video Restoration
- AgentAsk: Multi-Agent Systems Need to Ask
- Agent Bain vs. Agent McKinsey: A New Text-to-SQL Benchmark for the Business Domain
- L2M-AID: Autonomous Cyber-Physical Defense by Fusing Semantic Reasoning of Large Language Models with Multi-Agent Reinforcement Learning (Preprint)
- Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
- AMAS: Adaptively Determining Communication Topology for LLM-based Multi-Agent System
- Embracing Dialectic Intersubjectivity: Coordination of Differential Perspectives in Content Analysis With LLM Persona Simulation
- Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
- From Simulation to Strategy: Automating Personalized Interaction Planning for Conversational Agents
- Can Lessons From Human Teams Be Applied to Multi-Agent Systems? The Role of Structure, Diversity, and Interaction Dynamics
- ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems
- ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization
- LLMs as Policy-Agnostic Teammates: A Case Study in Human Proxy Design for Heterogeneous Agent Teams
- CAM: A Constructivist View of Agentic Memory for LLM-Based Reading Comprehension
- Bloom: Designing for LLM-Augmented Behavior Change Interactions
- MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
- QuantAgents: Towards Multi-agent Financial System via Simulated Trading
- GenQuest: An LLM-based Text Adventure Game for Language Learners
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
- Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- A Qualitative Comparative Evaluation of Cognitive and Generative Theories
- Prototyping Digital Social Spaces through Metaphor-Driven Design: Translating Spatial Concepts into an Interactive Social Simulation
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- RELATE-Sim: Leveraging Turning Point Theory and LLM Agents to Predict and Understand Long-Term Relationship Dynamics through Interactive Narrative Simulations
- Cyber Academia-Chemical Engineering (CA-ChemE): A Living Digital Town for Self-Directed Research Evolution and Emergent Scientific Discovery
- The Hunger Game Debate: On the Emergence of Over-Competition in Multi-Agent Systems
- Evaluating the Use of Large Language Models as Synthetic Social Agents in Social Science Research
- RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
- Leveraging LLMs to Improve Experimental Design: A Generative Stratification Approach
- Learning from Convenience Samples: A Case Study on Fine-Tuning LLMs for Survey Non-response in the German Longitudinal Election Study
- ID-RAG: Identity Retrieval-Augmented Generation for Long-Horizon Persona Coherence in Generative Agents
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- From Ambiguity to Verdict: A Semiotic-Grounded Multi-Perspective Agent for LLM Logical Reasoning
- Agentic Services Computing
- Memory Transfer Planning: LLM-driven Context-Aware Code Adaptation for Robot Manipulation
- PhysiAgent: An Embodied Agent Framework in Physical World
- Spiral of Silence in Large Language Model Agents
- Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark
- AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
- NeuroBridge: Using Generative AI to Bridge Cross-neurotype Communication Differences through Neurotypical Perspective-taking
- Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts
- Think Socially via Cognitive Reasoning
- The Emergence of Altruism in Large-Language-Model Agents Society
- Leveraging LLM Agents for Automated Video Game Testing
- Generalized Multi-agent Social Simulation Framework
- What Makes LLM Agent Simulations Useful for Policy? Insights From an Iterative Design Engagement in Emergency Preparedness
- Reimagining Agent-based Modeling with Large Language Model Agents via Shachi
- Synthetic Dialogue Generation for Interactive Conversational Elicitation & Recommendation (ICER)
- ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
- Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems
- Accelerate Creation of Product Claims Using Generative AI
- What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns
- SeMob: Semantic Synthesis for Dynamic Urban Mobility Prediction
- LLMs struggle to simulate human belief updates in controlled environments
- MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory
- Training Skills Like Parameters via Self-Supervised Semantic Diffusion
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games
- FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
- MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
- Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
- ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs
- ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory
- Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- AutoMem: Automated Learning of Memory as a Cognitive Skill
- Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates
- What If Prompt Injection Never Left? Rethinking Agent Security through Cross-Session Stored Prompt Injection
- Undoing Gracia: queering the self in the algorithmic borderlands
- AgentClinic: a multimodal benchmark for tool-using clinical AI agents
- StateScribe: Towards Accessible Change Awareness Across Real-World Revisits
- A Theory of Appropriateness That Accounts for Norms of Rationality
- Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and Considerations
- This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
- The Indispensable Role of User Simulation in the Pursuit of AGI
- Simulating Online Social Media Conversations on Controversial Topics Using AI Agents Calibrated on Real-World Data
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- Enhancing LLM-Based Social Bot via an Adversarial Learning Framework
- Agentic AutoSurvey: Let LLMs Survey LLMs
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- Bounded PCTL Model Checking of Large Language Model Outputs
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
- Human vs. Agent in Task-Oriented Conversations
- The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
- LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts
- OPEN-THEATRE: An Open-Source Toolkit for LLM-based Interactive Drama
- LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
- Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
- Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
- The Psychology of Falsehood: A Human-Centric Survey of Misinformation Detection
- Learning in Context: Personalizing Educational Content with Large Language Models to Enhance Student Learning
- ClearFairy: Capturing Creative Workflows through Decision Structuring, In-Situ Questioning, and Rationale Inference
- Foam-Agent 2.0: An End-to-End Composable Multi-Agent Framework for Automating CFD Simulation in OpenFOAM
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- When Avatars Have Personality: Effects on Engagement and Communication in Immersive Medical Training
- MICA: Multi-Agent Industrial Coordination Assistant
- Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
- AgentCTG: Harnessing Multi-Agent Collaboration for Fine-Grained Precise Control in Text Generation
- Inject, Fork, Compare: Defining an Interaction Vocabulary for Multi-Agent Simulation Platforms
- Programmable Cognitive Bias in Social Agents
- AI Agents with Human-Like Collaborative Tools: Adaptive Strategies for Enhanced Problem-Solving
- Evaluating LLM Alignment on Personality Inference from Real-World Interview Data
- A Visualized Framework for Event Cooperation with Generative Agents
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- Beyond Data Privacy: New Privacy Risks for Large Language Models
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
- DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow
- Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
- RAGs to Riches: RAG-like Few-shot Learning for Large Language Model Role-playing
- BioMetaphor: AI-Generated Biodata Representations for Virtual Co-Present Events
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble
- Evalet: Evaluating Large Language Models through Functional Fragmentation
- When Your Boss Is an AI Bot: Exploring Opportunities and Risks of Manager Clone Agents in the Future Workplace
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- RecoWorld: Building Simulated Environments for Agentic Recommender Systems
- Compartmentalised Agentic Reasoning for Clinical NLI
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- AI Wellbeing
- SEDM: Scalable Self-Evolving Distributed Memory for Agents
- Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
- Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining
- Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference
- Timing the Message: Language-Based Notifications for Time-Critical Assistive Settings
- Disentangling Interaction and Bias Effects in Opinion Dynamics of Large Language Models
- Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
- Simulating Dispute Mediation with LLM-Based Agents for Legal Research
- Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial
- Code2MCP: Transforming Code Repositories into MCP Services
- ProfilingAgent: Profiling-Guided Agentic Reasoning for Adaptive Model Optimization
- DualAlign: Generating Clinically Grounded Synthetic Data
- Internet 3.0: Architecture for a Web-of-Agents with it's Algorithm for Ranking Agents
- Emergent Social Dynamics of LLM Agents in the El Farol Bar Problem
- What Would an LLM Do? Evaluating Policymaking Capabilities of Large Language Models
- SasAgent: Multi-Agent AI System for Small-Angle Scattering Data Analysis
- Learning to Deliberate: Meta-policy Collaboration for Agentic LLMs with Multi-agent Reinforcement Learning
- Beyond Interpretability: Exploring the Comprehensibility of Adaptive Video Streaming through Large Language Models
- Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
- The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices
- Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
- RumorSphere: A Framework for Million-scale Agent-based Dynamic Simulation of Rumor Propagation
- Batch Query Processing and Optimization for Agentic Workflows
- Graph RAG as Human Choice Model: Building a Data-Driven Mobility Agent with Preference Chain
- Plantbot: Integrating Plant and Robot through LLM Modular Agent Networks
- Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
- Heads or Tails: A Simple Example of Causal Abstractive Simulation
- From CVE Entries to Verifiable Exploits: An Automated Multi-Agent Framework for Reproducing CVEs
- LLM-Assisted Iterative Evolution with Swarm Intelligence Toward SuperBrain
- Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
- The Resurgence of GCG Adversarial Attacks on Large Language Models
- Social World Models
- Synthetic Founders: AI-Generated Social Simulations for Startup Validation Research in Computational Social Science
- Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment
- COCORELI: Cooperative, Compositional Reconstitution & Execution of Language Instructions
- A Financial Brain Scan of the LLM
- ChatThero: An LLM-Supported Chatbot for Behavior Change and Therapeutic Support in Addiction Recovery
- Joint Enhancement of Relational Reasoning for Long-Context LLMs
- Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Instructional Agents: LLM Agents on Automated Course Material Generation for Teaching Faculties
- Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- DeepMEL: A Multi-Agent Collaboration Framework for Multimodal Entity Linking
- EMNLP: Educator-role Moral and Normative Large Language Models Profiling
- From Bits to Boardrooms: A Cutting-Edge Multi-Agent LLM Framework for Business Excellence
- A Concurrent Modular Agent: Framework for Autonomous LLM Agents
- MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
- STARec: An Efficient Agent Framework for Recommender Systems via Autonomous Deliberate Reasoning
- Bias-Adjusted LLM Agents for Human-Like Decision-Making via Behavioral Economics
- Open-Universe Assistance Games
- Can LLM Agents Solve Collaborative Tasks? A Study on Urgency-Aware Planning and Coordination
- From Passive Tool to Socio-cognitive Teammate: A Conceptual Framework for Agentic AI in Human-AI Collaborative Learning
- Virtual Community: An Open World for Humans, Robots, and Society
- LLM-Powered Virtual Patient Agents for Interactive Clinical Skills Training with Automated Feedback
- CausalPlan: Empowering Efficient LLM Multi-Agent Collaboration Through Causality-Driven Planning
- `My Dataset of Love': A Preliminary Mixed-Method Exploration of Human-AI Romantic Relationships
- Personas within Parameters: Fine-Tuning Small Language Models with Low-Rank Adapters to Mimic User Behaviors
- HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
- GTool: Graph Enhanced Tool Planning with Large Language Model
- Semantic Anchoring in Agentic Memory: Leveraging Linguistic Structures for Persistent Conversational Context
- Diagnostic-Guided Dynamic Profile Optimization for LLM-based User Simulators in Sequential Recommendation
- Do Large Language Model Agents Exhibit a Survival Instinct? An Empirical Study in a Sugarscape-Style Simulation
- Benchmarking LLM-based Agents for Single-cell Omics Analysis
- CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
- Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework
- MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance
- Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
- What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
- The PacifAIst Benchmark:Would an Artificial Intelligence Choose to Sacrifice Itself for Human Safety?
- Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
- The Roots of International Perceptions: Simulating US Attitude Changes Towards China with LLM Agents
- Simulating Generative Social Agents via Theory-Informed Workflow Design
- Exploring Large Language Model Agents for Piloting Social Experiments
- Cowpox: Towards the Immunity of VLM-based Multi-Agent Systems
- PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
- IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
- Livia: An Emotion-Aware AR Companion Powered by Modular AI Agents and Progressive Memory Compression
- DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
- Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
- Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
- MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
- FEAT: A Multi-Agent Forensic AI System with Domain-Adapted Large Language Model for Automated Cause-of-Death Analysis
- EMPATHIA: Multi-Faceted Human-AI Collaboration for Refugee Integration
- Agentic information systems
- Valid Inference with Imperfect Synthetic Data
- Scaling Personality Control in LLMs with Big Five Scaler Prompts
- LLMs for Resource Allocation: A Participatory Budgeting Approach to Inferring Preferences
- Simulating Human-Like Learning Dynamics with LLM-Empowered Agents
- The Term 'Agent' Has Been Diluted Beyond Utility and Requires Redefinition
- From MAS to MARS: Coordination Failures and Reasoning Trade-offs in Hierarchical Multi-Agent Robotic Systems within a Healthcare Scenario
- VirtLab: An AI-Powered System for Flexible, Customizable, and Large-scale Team Simulations
- DRAMA: A Dynamic and Robust Allocation-based Multi-Agent System for Changing Environments
- StackPilot: Autonomous Function Agents for Scalable and Environment-Free Code Execution
- Bridging Brains and Models: MoE-Based Functional Lesions for Simulating and Rehabilitating Aphasia
- ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"
- Galaxy: A Cognition-Centered Framework for Proactive, Privacy-Preserving, and Self-Evolving LLM Agents
- Sotopia-RL: Reward Design for Social Intelligence
- Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
- Parallelism Meets Adaptiveness: Scalable Documents Understanding in Multi-Agent LLM Systems
- Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- Who is a Better Player: LLM against LLM
- AgentSME for Simulating Diverse Communication Modes in Smart Education
- Tree-of-Reasoning: Towards Complex Medical Diagnosis via Multi-Agent Reasoning with Evidence Tree
- Survey of Large Language Models in Extended Reality: Technical Paradigms and Application Frontiers
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
- InqEduAgent: Adaptive AI Learning Partners with Gaussian Process Augmentation
- Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired
- A Multi-Agent System for Complex Reasoning in Radiology Visual Question Answering
- Trainable Dynamic Mask Sparse Attention
- CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search
- Whispering Agents: An Event-driven Covert Communication Protocol For the Internet of Agents
- Meta-RAG on Large Codebases Using Code Summarization
- Social Media Information Operations
- L3M+P: Lifelong Planning with Large Language Models
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- Applying Psychometrics to Large Language Model Simulated Populations: Recreating the HEXACO Personality Inventory Experiment with Generative Agents
- From Logic to Language: A Trust Index for Problem Solving with LLMs
- A survey of multi-agent geosimulation methodologies: from ABM to LLM
- LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra
- DynaSwarm: Dynamically Graph Structure Selection for LLM-based Multi-agent System
- A Framework for Analyzing Abnormal Emergence in Service Ecosystems Through LLM-based Agent Intention Mining
- GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts
- MetaAgent: Automatically Constructing Multi-Agent Systems Based on Finite State Machines
- Towards Simulating Social Influence Dynamics with LLM-based Multi-agents
- PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins
- Mitigating Response Delays in Free-Form Conversations with LLM-powered Intelligent Virtual Agents
- An Explainable Emotion Alignment Framework for LLM-Empowered Agent in Metaverse Service Ecosystem
- Strategic Communication and Language Bias in Multi-Agent LLM Coordination
- FlowForge: Guiding the Creation of Multi-agent Workflows with Design Space Visualization as a Thinking Scaffold
- Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel Predictions
- MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
- A Multi-Agent Generative AI Framework for IC Module-Level Verification Automation
- Evaluation and Benchmarking of LLM Agents: A Survey
- A2HCoder: An LLM-Driven Coding Agent for Hierarchical Algorithm-to-HDL Translation
- MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations
- ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable Devices
- Games Agents Play: Towards Transactional Analysis in LLM-based Multi-Agent Systems
- Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling
- MLC-Agent: Cognitive Model based on Memory-Learning Collaboration in LLM Empowered Agent Simulation Environment
- DynamiX: Large-Scale Dynamic Social Network Simulator
- Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
- CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback
- Integrating LLM in Agent-Based Social Simulation: Opportunities and Challenges
- Event-Driven Storytelling with Multiple Lifelike Humans in a 3D Scene
- PosterMate: Audience-driven Collaborative Persona Agents for Poster Design
- An Empirical Study of GenAI Adoption in Open-Source Game Development: Tools, Tasks, and Developer Challenges
- SYNTHIA: Synthetic Yet Naturally Tailored Human-Inspired PersonAs
- Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation
- A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
- Eywa: Provenance-Grounded Long-Term Memory for AI Agents
- Exploring the Dynamic Scheduling Space of Real-Time Generative AI Applications on Emerging Heterogeneous Systems
- A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents
- FutureSim: Replaying World Events to Evaluate Adaptive Agents
- SimPersona: Learning Discrete Buyer Personas from Raw Clickstreams for Grounded E-Commerce Agents
- Configurable multi-agent framework for scalable and realistic testing of llm-based agents
- Adaptive Multi-Agent Reasoning via Automated Workflow Generation
- Generative AI in Qualitative Research and Related Transparency Problems: A Novel Heuristic for Disclosing Uses of AI
- CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
- Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning
- Imitating Mistakes in a Learning Companion AI Agent for Online Peer Learning
- LLM Agent-Based Simulation of Student Activities and Mental Health Using Smartphone Sensing Data
- osmAG-LLM: Zero-Shot Open-Vocabulary Object Navigation via Semantic Maps and Large Language Models Reasoning
- Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
- Value-Based Large Language Model Agent Simulation for Mutual Evaluation of Trust and Interpersonal Closeness
- REVA: Supporting LLM-Generated Programming Feedback Validation at Scale Through User Attention-based Adaptation
- DiaryPlay: AI-Assisted Authoring of Interactive Vignettes for Everyday Storytelling
- Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
- Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
- Large Population Models
- TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit
- Negotiating Comfort: Simulating Personality-Driven LLM Agents in Shared Residential Social Networks
- AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs
- Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
- A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities
- Multi-Actor Generative Artificial Intelligence as a Game Engine
- StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
- KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows
- State-Inference-Based Prompting for Natural Language Trading with Game NPCs
- MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection
- A Mathematical Theory of Discursive Networks
- Agent-based model [wikipedia]
Discussions
- Generative Agents: Interactive Simulacra of Human Behavior [hn, 391 points, 252 comments]
- 1. A large and growing literature on "generative agent-based modeling" aims to study social science questions by using LLMs to simulate the behavior of people in naturalistic situations. I think it's [bsky, 193 points, 12 comments]
- Generative Agents: Interactive Simulacra of Human Behavior [hn, 13 points, 2 comments]
- New Advanced Video Game AI - Generative Agents: Interactive Simulacra of Human Behavior [lemmy, 13 points, 1 comments]
- Like, people have been working on this kind of problem for decades. Chris Crawford even galloped out of GDC on an imaginary horse and shunned the entire rest of the game industry for decades to work o [bsky, 7 points, 1 comments]
- אין לי כח לכל הדיבור על moltbook הרשת החברתית של הבוטים הנוראיים (בניגוד ל-X, הרשת החברתית של הבוטים הנוראיים). אבל לרקע - המאמר הזה מלפני שנתיים וחצי: arxiv.org/abs/2304.03442 [bsky, 6 points, 1 comments]
- My position is that agentic AI relies on “sharp edged” problems for which there are big costs if it makes a mistake, implying very high accuracy requirements—and for most agentic uses, we aren’t at th [bsky, 3 points, 1 comments]
- I have a few models in the works, but this one is inspired by this "Sims" proof of concept where LLM-powered agents interact quite freely with each other. Really interesting paper. As I said, just wat [bsky, 3 points, 1 comments]
- You can read the paper here: arxiv.org/pdf/2304.03442 [bsky, 3 points, 1 comments]
- [R] Generative Agents: Interactive Simulacra of Human Behavior (to be presented at UIST) [lemmy, 3 points, 0 comments]
- Generative Agents: Interactive Simulacra of Human Behavior [lobsters, 2 points, 0 comments]
- Oh ye, I remember I had to look into that for work. I think you are talking about this: arxiv.org/pdf/2304.03442? The outcome was pretty impressive, but there were some severe limitations in time of [bsky, 2 points, 1 comments]
- Fascinating stuff. I’m still reading through your article so perhaps you already mention this, but have you seen this paper? arxiv.org/pdf/2304.03442 [bsky, 1 points, 2 comments]
- This paper has a good and simple memory consolidation mechanism: arxiv.org/abs/2304.03442 [bsky, 1 points, 1 comments]
- In case you missed it, Stanford published a paper about AI agents: https://arxiv.org/pdf/2304.03442.pdf Generative agents wake up, cook breakfast, head to work; They form opinions, notice each other, [bsky, 1 points, 0 comments]
- This is the paper btw. arxiv.org/abs/2304.03442 I remember it because it was part of a seminar and looked interesting enough for the not so LLM skeptic me at first glance. How things have changed in t [bsky, 1 points, 0 comments]
- arxiv.org/abs/2304.03442 [bsky, 1 points, 0 comments]
- Stanford and Google researchers released 25 AI bots into a virtual town, Smallville. These agents cooked, worked, and socialized like humans. They roamed schools, cafés, and bars - like The Sims but w [bsky, 1 points, 0 comments]
- This is one of my favorite AI papers recently uses a network of LLMs to simulate a Stardew valley esque community. https://arxiv.org/abs/2304.03442 They observed some pretty incredible emergent soci [bsky, 0 points, 0 comments]
- @theophite.bsky.social do you know of any actual research backing this stuff or is it just utter nonsense? Lile they are partnering with this company Simile, and their only publication on their site i [bsky, 0 points, 0 comments]
- Check out this research paper from three years ago then. Even the earlier models were already pretty good at simulating basic human behavior with the right setup. arxiv.org/pdf/2304.03442 [bsky, 0 points, 0 comments]
- Generative Agents: Interactive Simulacra of Human Behavior arxiv.org/abs/2304.03442 [bsky, 0 points, 0 comments]
- I heard something similar a while ago,but it was not on Minecraft arxiv.org/abs/2304.034... It is really interesting I wonder if we would ever get NPC towns with this level of involvement [bsky, 0 points, 0 comments]
- Westworld vibes 🦄 https://arxiv.org/pdf/2304.03442.pdf [bsky, 0 points, 0 comments]
- At Stanford, a research team has created a city simulation wherein ChatGPT-trained “generative agents” appear to approximate “plausible” human behavior. Notably, they had “meaningful” conversations [bsky, 0 points, 1 comments]
Related