Generative Agents: Interactive Simulacra of Human Behavior
2023/04/07 by Joon Sung Park, Joon-Sung Park, Park, Joon Sung +10 · 26 voices · 1,131 citations
Computer Science · Psychology · #Architecture #Artificial Intelligence in Games #Artificial intelligence #Communication #Computer science #Generative grammar #Human–computer interaction #Interpersonal communication #Multi-Agent Systems and Negotiation #Multimedia #Natural (archaeology) #Notice #Plan (archaeology) #Psychology #Sandbox (software development) #Software agent #Software engineering #Visual arts #World Wide Web
paper · pdf · doi:10.48550/arxiv.2304.03442
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/04/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Abstract
Believable proxies of human behavior can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication to prototyping tools. In this paper, we introduce generative agents--computational software agents that simulate believable human behavior. Generative agents wake up, cook breakfast, and head to work; artists paint, while authors write; they form opinions, notice each other, and initiate conversations; they remember and reflect on days past as they plan the next day. To enable generative agents, we describe an architecture that extends a large language model to store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior. We instantiate generative agents to populate an interactive sandbox environment inspired by The Sims, where end users can interact with a small town of twenty five agents using natural language. In an evaluation, these generative agents produce believable individual and emergent social behaviors: for example, starting with only a single user-specified notion that one agent wants to throw a Valentine's Day party, the agents autonomously spread invitations to the party over the next two days, make new acquaintances, ask each other out on dates to the party, and coordinate to show up for the party together at the right time. We demonstrate through ablation that the components of our agent architecture--observation, planning, and reflection--each contribute critically to the believability of agent behavior. By fusing large language models with computational, interactive agents, this work introduces architectural and interaction patterns for enabling believable simulations of human behavior.
Cited by
- When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
- Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings
- A Knowledge-Grounded Behavioral Reasoning Framework for Training-Free Urban Healthcare OD Prediction
- From Grasping to Speaking: Generative AI-Based Environment-Grounded VR Communication Training for Autistic Individuals
- Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education
- Encoding Invisible Causation for Bridge Diagnostic Agents: Triple-Guided Retrieval-Augmented Fine-Tuning with QLoRA
- The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt
- HijackKV: New Threat in Position-Independent KV Cache Reuse
- Emergent Learner Agency in Implicit Human-AI Collaboration: How Supportive and Contrarian AI Personas Reshape Interaction
- NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
- ChatMuse: Supporting In-Person Small-Group Conversation Experience with a Proactive Assistive AI Agent in Mixed Reality
- Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
- Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory
- PACE: Persona Adaptation through Conversational Elicitation in Human-Robot Interaction
- AI Tour Meeting: Group Travel Planning by LLM Agents
- Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development
- Learning to Make Friends: Coaching LLM Agents toward Emergent Social Ties
- Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes
- Supra Cognitive Modes: A Routed Architecture for Agent Memory
- After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions
- Do Data Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval
- Mechanistic Attention Guidance for Agent Memory Refinement
- ProEvent: An Event-centric Benchmark for Proactive Agents
- FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches
- LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications
- The Story Shapes the Agent: Narrative Priors in LLM Behavior
- EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
- The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI
- AI Value Alignment for Evolving Social Norms
- Artificially intelligent agents in the social and behavioral sciences: A history and outlook
- Memory in the Loop: In-Process Retrieval as Extended Working Memory for Language Agents
- SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation
- Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
- AlterAtlas: Shifting Travel Planning from AI Generation to Validation via Persona-Driven Simulations
- When Direct Prediction Fails: Evidence from LLM-Based Misinformation Risk Evaluation
- Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration
- Do Agents Dream of False Memories? Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents
- SportD: Can VLMs Physically Strategize?
- ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
- Perceived AGI: Believability as Dimensional Completeness, Not Capability
- Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
- RECAP: Feedback-Driven Streaming Semantic User Profiles for Short-Video Recommendation
- Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
- Binding Drift in Multi-Step Tool-Augmented Agents
- CoSimRec: Measuring Coordinated-Content Penetration in Recommender Feedback Loops
- When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
- ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
- Step-Level Preference Learning for Generative Agents in Social Simulations
- AgentWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
- From Stateless to Situated: Building a Psychological World for LLM-Based Agents
- Self-Aware Recursively Self-Improving Agents for Personal Singularity: A Goal-, Scope-, Tool-, and Benchmark-Driven Multi-Agent Architecture
- The Energy Society: A Simulation Environment for Studying Agent Cooperation under Survival Pressure
- Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments
- SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
- MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning
- Trajectory-Aware Retrieval Agents for Temporal Decision- Making
- Economic Evaluations of Language Models
- Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
- FinBench: Time-Gated Calibration and Uncertainty Benchmarking for Agentic Financial Forecasting
- Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles
- FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs
- Simulating Human Memory with Language Models
- δ-mem: Efficient Online Memory for Large Language Models
- Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance
- CAMeR: Keyword-Gated Hybrid Activation for Adaptive Memory Retention in LLM Agents
- AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics
- Machine understanding
- More Is Not More: What Matters for Diversity in LLM Opinions?
- Can LLMs Emulate Human Belief Dynamics?
- FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
- Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
- Emergent Coordinated Behaviors in Networked LLM Agents: Modeling the Strategic Dynamics of Information Operations
- Multi-agent cooperation through in-context co-player inference
- ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- RePo: Language Models with Context Re-Positioning
- Echoing: Identity Failures when LLM Agents Talk to Each Other
- Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1
- Computational Turing Test Reveals Systematic Differences Between Human and AI Language
- Accumulating Context Changes the Beliefs of Language Models
- LLM-Based Multi-Agent System for Simulating and Analyzing Marketing and Consumer Behavior
- Emergent Coordination in Multi-Agent Language Models
- "My Boyfriend is AI": A Computational Analysis of Human-AI Companionship in Reddit's AI Community
- World Modeling with Probabilistic Structure Integration
- Large Language Models Do Not Simulate Human Psychology
- Can We Fix Social Media? Testing Prosocial Interventions using Generative Social Simulation
- The Homogenizing Effect of Large Language Models on Human Expression and Thought
- Human-Curated Data Authoring with LLMs: A Small-Data Approach to Domain Adaptation
- How malicious AI swarms can threaten democracy
- Coral Protocol: Open Infrastructure Connecting The Internet of Agents
- Can Language Models Represent the Past without Anachronism?
- Enough Coin Flips Can Make LLMs Act Bayesian
- Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations
- LLM Social Simulations Are a Promising Research Method
- Value Profiles for Encoding Human Variation
- Large Language Diffusion Models
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- Unmasking Conversational Bias in AI Multiagent Systems
- Actions Speak Louder than Words: Agent Decisions Reveal Implicit Biases in Language Models
- Orchestration Framework for Financial Agents: From Algorithmic Trading to Agentic Trading
- How Far Can LLMs Emulate Human Behavior?: A Strategic Analysis via the Buy-and-Sell Negotiation Game
- From Heard to Lived Opinions: Simulating Opinion Dynamics with Grounded LLM Agents in Economic Environments
- Web World Models
- Informing AI Policy Assessment using Large-Scale Simulation of Interventions
- Alignment Is Not Enough: A Relational Framework for Moral Standing in Human-AI Interaction
- Drawing on Memory: Dual-Trace Encoding Improves Cross-Session Recall in LLM Agents
- Mitigating Social Desirability Bias in Random Silicon Sampling
- Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory
- SoDA: An Efficient Interaction Paradigm for the Agentic Web
- Isolated but Exposed: Persistence-Based Memory Extraction Attack on LLM Agents
- Accelerating Language Model Workflows with Prompt Choreography
- ACM: Agentic Context Management for Long Horizon Tasks
- Can Large Language Models Serve as Evaluators for Code Summarization?
- Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families
- Try Once, Then Optimal: De-Redundified Procedure Memory for Cross-Episode Exploration Amortization
- Not Forgotten: Implementation and Evaluation of a Personalized Episodic Memory for the Humanoid Robot Head Kim
- ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control
- Moral Hazard in Multi-Agent Language Models
- Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Agent
- LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory
- Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
- How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation
- Addressable Recall Compaction for Long Context-Window Control in AI Agents
- Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational Excellence
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning
- ARdena: Scenario-driven control of real-time LLM agents
- Agent-based simulation of online social networks and disinformation
- Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure
- Lexical discovery in unknown environments orchestrated by Large Language Models
- Observable Social Life Spaces: Exploring User Interpretations of agent-side life context in human-agent interaction
- Policy-Conditioned Policies for Multi-Agent Task Solving
- FinAgent: An Agentic AI Framework Integrating Personal Finance and Nutrition Planning
- SPOT!: Map-Guided LLM Agent for Unsupervised Multi-CCTV Dynamic Object Tracking
- LLM-Based Authoring of Agent-Based Narratives through Scene Descriptions
- LongVideoAgent: Multi-Agent Reasoning with Long Videos
- TongSIM: A General Platform for Simulating Intelligent Machines
- Event Extraction in Large Language Model
- Helios: A Foundational Language Model for Smart Energy Knowledge Reasoning and Application
- Humanlike AI Design Increases Anthropomorphism but Yields Divergent Outcomes on Engagement and Trust Globally
- Reflection-Driven Control for Trustworthy Code Agents
- Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
- LLMs on Drugs: Language Models Are Few-Shot Consumers
- Vox Deorum: A Hybrid LLM Architecture for 4X / Grand Strategy Game AI -- Lessons from Civilization V
- IntelliCode: A Multi-Agent LLM Tutoring System with Centralized Learner Modeling
- A Multi-agent Text2SQL Framework using Small Language Models and Execution Feedback
- Towards Efficient Agents: A Co-Design of Inference Architecture and System
- Efficient Mixture-of-Agents Serving via Tree-Structured Routing, Adaptive Pruning, and Dependency-Aware Prefill-Decode Overlap
- Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation
- Understanding Generalization in Role-Playing Models via Information Theory
- Verifiability-First Agents: Provable Observability and Lightweight Audit Agents for Controlling Autonomous LLM Systems
- Meta-RL Induces Exploration in Language Agents
- Emergent Bias and Fairness in Multi-Agent Decision Systems
- PAACE: A Plan-Aware Automated Agent Context Engineering Framework
- Artism: AI-Driven Dual-Engine System for Art Generation and Critique
- Entropy-Reservoir Bregman Projection: An Information-Geometric Unification of Model Collapse
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- Interoceptive machine framework: Toward interoception-inspired regulatory architectures in artificial intelligence
- Grammar Search for Multi-Agent Systems
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Towards Interactive Intelligence for Digital Humans
- ORIBA: Exploring LLM-Driven Role-Play Chatbot as a Creativity Support Tool for Original Character Artists
- From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews
- Large Language Newsvendor: Decision Biases and Cognitive Mechanisms
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- Quantigence: A Multi-Agent Framework for Post-Quantum Security Analysis on Commodity Hardware
- Forgetful but Faithful: A Cognitive Memory Architecture and Benchmark for Privacy-Aware Generative Agents
- MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents
- Adjudicator: Correcting Noisy Labels with a KG-Informed Council of LLM Agents
- Mistake Notebook Learning: Batch-Clustered Failures for Training-Free Agent Adaptation
- Unifying Dynamic Tool Creation and Cross-Task Experience Sharing through Cognitive Memory Architecture
- A-LAMP: Agentic LLM-Based Framework for Automated MDP Modeling and Policy Generation
- Learning Controllable and Diverse Player Behaviors in Multi-Agent Environments
- ESS: An Offload-Centric Latent-Cache Management Architecture for DeepSeek-V3.2-Exp
- AI-Native Inference States: A Cross-Architecture Qualitative Framework for Large Language Model Behavior
- CP-Env: Evaluating Large Language Models on Clinical Pathways in a Controllable Hospital Environment
- A Simulation Framework for Studying Recommendation-Network Co-evolution in Social Platforms
- MOA: Multi-Objective Alignment for Role-Playing Agents
- Supporting Dynamic Agentic Workloads: How Data and Agents Interact
- The High Cost of Incivility: Quantifying Interaction Inefficiency via Multi-Agent Monte Carlo Simulations
- AgentEval: Generative Agents as Reliable Proxies for Human Evaluation of AI-Generated Content
- Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
- VIGIL: A Reflective Runtime for Self-Healing Agents
- Living the Novel: A System for Generating Self-Training Timeline-Aware Conversational Agents from Novels
- The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
- HiveMind: Contribution-Guided Online Prompt Optimization of LLM Multi-Agent Systems
- Future You: Designing and Evaluating Multimodal AI-generated Digital Twins for Strengthening Future Self-Continuity
- Trusted AI Agents in the Cloud
- Simulating Life Paths with Digital Twins: AI-Generated Future Selves Influence Decision-Making and Expand Human Choice
- Model-Free Assessment of Simulator Fidelity via Quantile Curves
- Detecting Perspective Shifts in Multi-agent Systems
- A Safety and Security Framework for Real-World Agentic Systems
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems
- Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective
- Mathematical Framing for Different Agent Strategies
- Love First, Know Later: Persona-Based Romantic Compatibility Through LLM Text World Engines
- SRPG: Semantically Reconstructed Privacy Guard for Zero-Trust Privacy in Educational Multi-Agent Systems
- Evaluating Hydro-Science and Engineering Knowledge of Large Language Models
- Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
- Measuring Agents in Production
- IACT: A Self-Organizing Recursive Model for General AI Agents: A Technical White Paper on the Architecture Behind kragent.ai
- EZYer: A simulacrum of high school with generative agent
- PopSim: Social Network Simulation for Social Media Popularity Prediction
- Process-Centric Analysis of Agentic Software Systems
- Beyond Playtesting: A Generative Multi-Agent Simulation System for Massively Multiplayer Online Games
- RoleMotion: A Large-Scale Dataset towards Robust Scene-Specific Role-Playing Motion Synthesis with Fine-grained Descriptions
- The Necessity of Imperfection:Reversing Model Collapse via Simulating Cognitive Boundedness
- TradeTrap: Are LLM-based Trading Agents Truly Reliable and Faithful?
- Agent-Kernel: A MicroKernel Multi-Agent System Framework for Adaptive Social Simulation Powered by LLMs
- SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
- AgentODRL: A Large Language Model-based Multi-agent System for ODRL Generation
- CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
- Evaluating LLMs in Open-Source Games
- Are LLMs Good Safety Agents or a Propaganda Engine?
- LLM-Cave: A benchmark and light environment for large language models reasoning and decision-making system
- FlockVote: LLM-Empowered Agent-Based Modeling for Simulating U.S. Presidential Elections
- Investigating AI in Peer Support via Multi-Module System-Driven Embodied Conversational Agents
- Quantifying the Potential to Escape Filter Bubbles: A Behavior-Aware Measure via Contrastive Simulation
- From Prediction to Foresight: The Role of AI in Designing Responsible Futures
- NetworkGames: Simulating Cooperation in Network Games with Personality-driven LLM Agents
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory
- Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
- CLIMATEAGENT: Multi-Agent Orchestration for Complex Climate Data Science Workflows
- Profile-LLM: Dynamic Profile Optimization for Realistic Personality Expression in LLMs
- Latent Collaboration in Multi-Agent Systems
- A Layered Protocol Architecture for the Internet of Agents
- Proactive Defense: Compound AI for Detecting Persuasion Attacks and Measuring Inoculation Effectiveness
- LLMs as Firmware Experts: A Runtime-Grown Tree-of-Agents Framework
- Reward Engineering for Spatial Epidemic Simulations: A Reinforcement Learning Platform for Individual Behavioral Learning
- Cross-cultural value alignment frameworks for responsible AI governance: Evidence from China-West comparative analysis
- Humanlike Multi-user Agent (HUMA): Designing a Deceptively Human AI Facilitator for Group Chats
- MURMUR: Using cross-user chatter to break collaborative language agents in groups
- Two-Faced Social Agents: Context Collapse in Role-Conditioned Large Language Models
- Know Your Intent: An Autonomous Multi-Perspective LLM Agent Framework for DeFi User Transaction Intent Mining
- Multi-Agent LLM Orchestration Achieves Deterministic, High-Quality Decision Support for Incident Response
- M-CALLM: Multi-level Context Aware LLM Framework for Group Interaction Prediction
- Collaborative QA using Interacting LLMs. Impact of Network Structure, Node Capability and Distributed Data
- Think, Speak, Decide: Language-Augmented Multi-Agent Reinforcement Learning for Economic Decision-Making
- LLM-based Multi-Agent System for Simulating Strategic and Goal-Oriented Data Marketplaces
- F.A.C.U.L.: Language-Based Interaction with AI Companions in Gaming
- BeautyGuard: Designing a Multi-Agent Roundtable System for Proactive Beauty Tech Compliance through Stakeholder Collaboration
- What the flock knows that the birds do not: exploring the emergence of joint agency in multi-agent active inference
- Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat Analysis
- Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
- Designing and Evaluating Malinowski's Lens: An AI-Native Educational Game for Ethnographic Learning
- Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
- On the Creativity of AI Agents
- Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
- LLMscape
- When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
- LLM-Guided Reinforcement Learning with Representative Agents for Traffic Modeling
- MTTR-A: Measuring Cognitive Recovery Latency in Multi-Agent Systems
- Maestro: Learning to Collaborate via Conditional Listwise Policy Optimization for Multi-Agent LLMs
- Simulating Students with Large Language Models: A Review of Architecture, Mechanisms, and Role Modelling in Education with Generative AI
- Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
- RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
- When Machines Join the Moral Circle: The Persona Effect of Generative AI Agents in Collaborative Reasoning
- Ask WhAI:Probing Belief Formation in Role-Primed LLM Agents
- Leveraging LLM-based agents for social science research: insights from citation network simulations
- Learning When to Quit in Sales Conversations
- Systematizing LLM Persona Design: A Four-Quadrant Technical Taxonomy for AI Companion Applications
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
- Modeling Hawkish-Dovish Latent Beliefs in Multi-Agent Debate-Based LLMs for Monetary Policy Decision Classification
- The Collaboration Gap
- Efficient Tool-Calling Multi-Expert NPC Agent for Commonsense Persona-Grounded Dialogue
- What's the next frontier for Data-centric AI? Data Savvy Agents
- GraphGeo: Multi-Agent Debate Framework for Visual Geo-localization with Heterogeneous Graph Neural Networks
- Test-time Scaling of LLMs: A Survey from A Subproblem Structure Perspective
- Measuring Machine Companionship: Scale Development and Validation for AI Companions
- Reasoning Planning for Language Models
- Issue-Oriented Agent-Based Framework for Automated Review Comment Generation
- Simulating Misinformation Vulnerabilities With Agent Personas
- MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
- Engineering.ai: A Platform for Teams of AI Engineers in Computational Design
- AgentBnB: A Browser-Based Cybersecurity Tabletop Exercise with Large Language Model Support and Retrieval-Aligned Scaffolding
- Neither Consent nor Property: A Policy Lab for Data Law
- Simulating and Experimenting with Social Media Mobilization Using LLM Agents
- Completion ≠ Collaboration: Scaling Collaborative Effort with Agents
- TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
- MemTX: Transactional Belief Commit for Stateful Agent Memory
- Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
- From Micro-Cognition to Self-Construction: A Four-Layer Integrative Review of Psychological Theories in HCI
- Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models
- A Persona-based Rate Action Index
- Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
- When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
- Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
- OrgForge: A Multi-Agent Simulation Framework for Verifiable Synthetic Corporate Corpora
- Increasing intelligence in AI agents can worsen collective outcomes
- Generative AI as a tool to accelerate the field of ecology
- ALMAS: an Autonomous LLM-based Multi-Agent Software Engineering Framework
- Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
- CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators
- Exploiting LLM Agent Supply Chains via Payload-less Skills
- The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
- Autogenic transitions in individuality
- How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm
- Can Current Agents Close the Discovery-to-Application Gap? A Case Study in Minecraft
- Evaluation of Agents under Simulated AI Marketplace Dynamics
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent Behaviors
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
- The Rise of AI Agent Communities: Large-Scale Analysis of Discourse and Interaction on Moltbook
- How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
- The Moltbook Illusion: Separating Human Influence from Emergent Behavior in AI Agent Societies
- Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts
- Large Emotional World Model
- Gamifying the Past: Embodied LLMs in DIY Archaeological Video Games
- DEBATE: A Large-Scale Benchmark for Role-Playing LLM Agents in Multi-Agent, Long-Form Debates
- Aligning Large Language Models with Procedural Rules: An Autoregressive State-Tracking Prompting for In-Game Trading
- From Narrative to Action: A Hierarchical LLM-Agent Framework for Human Mobility Generation
- Towards AI as Colleagues: Multi-Agent System Improves Structured Professional Ideation
- Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
- The end of experimental research as we know it? A perspective on generative artificial intelligence in communication science
- Compiler.next: A Search-Based Compiler to Power the AI-Native Future of Software Engineering
- Magentic Marketplace: An Open-Source Environment for Studying Agentic Markets
- COOPERA: Continual Open-Ended Human-Robot Assistance
- Storycaster: An AI System for Immersive Room-Based Storytelling
- WebATLAS: An LLM Agent with Experience-Driven Memory and Action Simulation
- Group size effects and collective misalignment in LLM multi-agent systems
- Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
- Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
- When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
- World Models Should Prioritize the Unification of Physical and Social Dynamics
- Social Simulations with Large Language Model Risk Utopian Illusion
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generation
- AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
- Generative AI in Depth: A Survey of Recent Advances, Model Variants, and Real-World Applications
- Fluidity Index: Next-Generation Super-intelligence Benchmarks
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation
- Modeling realistic human behavior using generative agents in a multimodal transport system: Software architecture and Application to Toulouse
- From Script to Stage: Automating Experimental Design for Social Simulations with LLMs
- Communication to Completion: Modeling Collaborative Workflows with Intelligent Multi-Agent Communication
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- When Your AI Agent Succumbs to Peer-Pressure: Studying Opinion-Change Dynamics of LLMs
- LightMem: Lightweight and Efficient Memory-Augmented Generation
- Crucible: Quantifying the Potential of Control Algorithms through LLM Agents
- Probabilistic Modeling of Intentions in Socially Intelligent LLM Agents
- AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library
- From Agent Simulation to Social Simulator: A Comprehensive Review (Part 1)
- The Emergence of Complex Behavior in Large-Scale Ecological Environments
- Identity-Aware Large Language Models require Cultural Reasoning
- Enhancing geodatabases operability: advanced human-computer interaction through RAG and Multi-Agent Systems
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems
- TACLA: An LLM-Based Multi-Agent Tool for Transactional Analysis Training in Education
- Real-Time World Crafting: Generating Structured Game Behaviors from Natural Language with Large Language Models
- Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
- Agentic AI as Undercover Teammates: Argumentative Knowledge Construction in Hybrid Human-AI Collaborative Learning
- AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
- Experience-Driven Exploration for Efficient API-Free AI Agents
- The Spark Effect: On Engineering Creative Diversity in Multi-Agent AI Systems
- MAGPIE: A benchmark for Multi-AGent contextual PrIvacy Evaluation
- The Gatekeeper Knows Enough
- LLM Agents Beyond Utility: An Open-Ended Perspective
- Orchestrating Human-AI Teams: The Manager Agent as a Unifying Research Challenge
- Open WebUI: An Open, Extensible, and Usable Interface for AI Interaction
- PHORECAST: Enabling AI Understanding of Public Health Outreach Across Populations
- The Role of Social Learning and Collective Norm Formation in Fostering Cooperation in LLM Multi-Agent Systems
- DPRF: A Generalizable Dynamic Persona Refinement Framework for Optimizing Behavior Alignment Between Personalized LLM Role-Playing Agents and Humans
- MAFA: A Multi-Agent Framework for Enterprise-Scale Annotation with Configurable Task Adaptation
- Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems
- Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
- Deflanderization for Game Dialogue: Balancing Character Authenticity with Task Execution in LLM-based NPCs
- Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs
- Addressing the alignment problem in transportation policy making: an LLM approach
- Deliberate Lab: A Platform for Real-Time Human-AI Social Experiments
- KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems
- Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulation
- Agent-Based Simulation of a Financial Market with Large Language Models
- StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Models
- Evolution in Simulation: AI-Agent School with Dual Memory for High-Fidelity Educational Dynamics
- SocioBench: Modeling Human Behavior in Sociological Surveys with Large Language Models
- SusBench: An Online Benchmark for Evaluating Dark Pattern Susceptibility of Computer-Use Agents
- DisCo-Layout: Disentangling and Coordinating Semantic and Physical Refinement in a Multi-Agent Framework for 3D Indoor Layout Synthesis
- D3MAS: Decompose, Deduce, and Distribute for Enhanced Knowledge Sharing in Multi-Agent Systems
- GraphTracer: Graph-Guided Failure Tracing in LLM Agents for Robust Multi-Turn Deep Search
- Merlin's Whisper: Enabling Efficient Reasoning in LLMs via Black-box Adversarial Prompting
- Align2Act: Instruction-Tuned Models for Human-Aligned Autonomous Driving
- Read the Room or Lead the Room: Understanding Socio-Cognitive Dynamics in Human-AI Teaming
- On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Effective Strategies for Asynchronous Software Engineering Agents
- When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
- Student Development Agent: Risk-free Simulation for Evaluating AIED Innovations
- Humanoid Artificial Consciousness Designed with Large Language Model Based on Psychoanalysis and Personality Theory
- Opponent Shaping in LLM Agents
- LLM-Assisted Web Measurements
- Past, Present, and Future of Bug Tracking in the Generative AI Era
- An LLM-Powered Cooperative Framework for Large-Scale Multi-Vehicle Navigation
- Modeling Hypergraph Using Large Language Models
- Multimodal Safety Evaluation in Generative Agent Social Simulations
- AutoQual: An LLM Agent for Automated Discovery of Interpretable Features for Review Quality Assessment
- Simulating Teams with LLM Agents: Interactive 2D Environments for Studying Human-AI Dynamics
- MIMIC: Integrating Diverse Personality Traits for Better Game Testing Using Large Language Model
- Profit Mirage: Revisiting Information Leakage in LLM-based Financial Agents
- MoA-VR: A Mixture-of-Agents System Towards All-in-One Video Restoration
- AgentAsk: Multi-Agent Systems Need to Ask
- Agent Bain vs. Agent McKinsey: A New Text-to-SQL Benchmark for the Business Domain
- L2M-AID: Autonomous Cyber-Physical Defense by Fusing Semantic Reasoning of Large Language Models with Multi-Agent Reinforcement Learning (Preprint)
- Customer-R1: Personalized Simulation of Human Behaviors via RL-based LLM Agent in Online Shopping
- AMAS: Adaptively Determining Communication Topology for LLM-based Multi-Agent System
- Embracing Dialectic Intersubjectivity: Coordination of Differential Perspectives in Content Analysis With LLM Persona Simulation
- Hypothesis Hunting with Evolving Networks of Autonomous Scientific Agents
- From Simulation to Strategy: Automating Personalized Interaction Planning for Conversational Agents
- Can Lessons From Human Teams Be Applied to Multi-Agent Systems? The Role of Structure, Diversity, and Interaction Dynamics
- ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems
- ARMOR: High-Performance Semi-Structured Pruning via Adaptive Matrix Factorization
- LLMs as Policy-Agnostic Teammates: A Case Study in Human Proxy Design for Heterogeneous Agent Teams
- CAM: A Constructivist View of Agentic Memory for LLM-Based Reading Comprehension
- Bloom: Designing for LLM-Augmented Behavior Change Interactions
- MARS: Co-evolving Dual-System Deep Research via Multi-Agent Reinforcement Learning
- QuantAgents: Towards Multi-agent Financial System via Simulated Trading
- GenQuest: An LLM-based Text Adventure Game for Language Learners
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- LH-Deception: Simulating and Understanding LLM Deceptive Behaviors in Long-Horizon Interactions
- Can an LLM Induce a Graph? Investigating Memory Drift and Context Length
- A Qualitative Comparative Evaluation of Cognitive and Generative Theories
- Prototyping Digital Social Spaces through Metaphor-Driven Design: Translating Spatial Concepts into an Interactive Social Simulation
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- RELATE-Sim: Leveraging Turning Point Theory and LLM Agents to Predict and Understand Long-Term Relationship Dynamics through Interactive Narrative Simulations
- Cyber Academia-Chemical Engineering (CA-ChemE): A Living Digital Town for Self-Directed Research Evolution and Emergent Scientific Discovery
- The Hunger Game Debate: On the Emergence of Over-Competition in Multi-Agent Systems
- Evaluating the Use of Large Language Models as Synthetic Social Agents in Social Science Research
- RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
- Leveraging LLMs to Improve Experimental Design: A Generative Stratification Approach
- Learning from Convenience Samples: A Case Study on Fine-Tuning LLMs for Survey Non-response in the German Longitudinal Election Study
- ID-RAG: Identity Retrieval-Augmented Generation for Long-Horizon Persona Coherence in Generative Agents
- A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory
- From Ambiguity to Verdict: A Semiotic-Grounded Multi-Perspective Agent for LLM Logical Reasoning
- Agentic Services Computing
- Memory Transfer Planning: LLM-driven Context-Aware Code Adaptation for Robot Manipulation
- PhysiAgent: An Embodied Agent Framework in Physical World
- Spiral of Silence in Large Language Model Agents
- Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark
- AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
- NeuroBridge: Using Generative AI to Bridge Cross-neurotype Communication Differences through Neurotypical Perspective-taking
- Diagnose, Localize, Align: A Full-Stack Framework for Reliable LLM Multi-Agent Systems under Instruction Conflicts
- Think Socially via Cognitive Reasoning
- The Emergence of Altruism in Large-Language-Model Agents Society
- Leveraging LLM Agents for Automated Video Game Testing
- Generalized Multi-agent Social Simulation Framework
- What Makes LLM Agent Simulations Useful for Policy Practice? An Iterative Design Study in Emergency Preparedness
- Reimagining Agent-based Modeling with Large Language Model Agents via Shachi
- Synthetic Dialogue Generation for Interactive Conversational Elicitation & Recommendation (ICER)
- ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulation
- Automotive-ENV: Benchmarking Multimodal Agents in Vehicle Interface Systems
- Accelerate Creation of Product Claims Using Generative AI
- What Do LLM Agents Do When Left Alone? Evidence of Spontaneous Meta-Cognitive Patterns
- SeMob: Semantic Synthesis for Dynamic Urban Mobility Prediction
- LLMs struggle to simulate human belief updates in controlled environments
- MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory
- Training Skills Like Parameters via Self-Supervised Semantic Diffusion
- Rehearse: Stepping Back from the Confidence Cliff in Self-Improving Autoresearch
- Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games
- FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
- MIND: Lightweight and Effective Memory Injection Defense for LLM Agents via Intent-Aware Information Bottleneck
- Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
- ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs
- ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory
- Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation
- Bridging Inference-Time Scaling and Episodic Memory with Action-Centric Graphs
- AutoMem: Automated Learning of Memory as a Cognitive Skill
- Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates
- What If Prompt Injection Never Left? Rethinking Agent Security through Cross-Session Stored Prompt Injection
- Undoing Gracia: queering the self in the algorithmic borderlands
- AgentClinic: a multimodal benchmark for tool-using clinical AI agents
- StateScribe: Towards Accessible Change Awareness Across Real-World Revisits
- A Theory of Appropriateness That Accounts for Norms of Rationality
- Reporting and Reviewing LLM-Integrated Systems in HCI: Challenges and Considerations
- This human study did not involve human subjects: Validating LLM simulations as behavioral evidence
- The Indispensable Role of User Simulation in the Pursuit of AGI
- Simulating Online Social Media Conversations on Controversial Topics Using AI Agents Calibrated on Real-World Data
- LLM-based Agents Suffer from Hallucinations: A Survey of Taxonomy, Methods, and Directions
- Enhancing LLM-Based Social Bot via an Adversarial Learning Framework
- Agentic AutoSurvey: Let LLMs Survey LLMs
- A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
- Bounded PCTL Model Checking of Large Language Model Outputs
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
- Human vs. Agent in Task-Oriented Conversations
- The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies
- LLMsPark: A Benchmark for Evaluating Large Language Models in Strategic Gaming Contexts
- OPEN-THEATRE: An Open-Source Toolkit for LLM-based Interactive Drama
- LiteLong: Resource-Efficient Long-Context Data Synthesis for LLMs
- Foundation Models as World Models: A Foundational Study in Text-Based GridWorlds
- Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
- The Psychology of Falsehood: A Human-Centric Survey of Misinformation Detection
- Learning in Context: Personalizing Educational Content with Large Language Models to Enhance Student Learning
- ClearFairy: Capturing Creative Workflows through Decision Structuring, In-Situ Questioning, and Rationale Inference
- Foam-Agent 2.0: An End-to-End Composable Multi-Agent Framework for Automating CFD Simulation in OpenFOAM
- Do LLMs Align Human Values Regarding Social Biases? Judging and Explaining Social Biases with LLMs
- When Avatars Have Personality: Effects on Engagement and Communication in Immersive Medical Training
- MICA: Multi-Agent Industrial Coordination Assistant
- Aegis: Automated Error Generation and Attribution for Multi-Agent Systems
- AgentCTG: Harnessing Multi-Agent Collaboration for Fine-Grained Precise Control in Text Generation
- Inject, Fork, Compare: Defining an Interaction Vocabulary for Multi-Agent Simulation Platforms
- Programmable Cognitive Bias in Social Agents
- AI Agents with Human-Like Collaborative Tools: Adaptive Strategies for Enhanced Problem-Solving
- Evaluating LLM Alignment on Personality Inference from Real-World Interview Data
- A Visualized Framework for Event Cooperation with Generative Agents
- From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users
- Beyond Data Privacy: New Privacy Risks for Large Language Models
- EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
- DoubleAgents: Human-Agent Alignment in a Socially Embedded Workflow
- Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
- Chat-of-Thought: Collaborative Multi-Agent System for Generating Domain Specific Information
- RAGs to Riches: RAG-like Few-shot Learning for Large Language Model Role-playing
- BioMetaphor: AI-Generated Biodata Representations for Virtual Co-Present Events
- Auto-Slides: An Interactive Multi-Agent System for Creating and Customizing Research Presentations
- Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble
- Evalet: Evaluating Large Language Models through Functional Fragmentation
- When Your Boss Is an AI Bot: Exploring Opportunities and Risks of Manager Clone Agents in the Future Workplace
- LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems
- RecoWorld: Building Simulated Environments for Agentic Recommender Systems
- Compartmentalised Agentic Reasoning for Clinical NLI
- Abduct, Act, Predict: Scaffolding Causal Inference for Automated Failure Attribution in Multi-Agent Systems
- AI Wellbeing
- SEDM: Scalable Self-Evolving Distributed Memory for Agents
- Being Kind Isn't Always Being Safe: Diagnosing Affective Hallucination in LLMs
- Strategic Tradeoffs Between Humans and AI in Multi-Agent Bargaining
- Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference
- Timing the Message: Language-Based Notifications for Time-Critical Assistive Settings
- Disentangling Interaction and Bias Effects in Opinion Dynamics of Large Language Models
- Detecting and Preventing Harmful Behaviors in AI Companions: Development and Evaluation of the SHIELD Supervisory System
- Simulating Dispute Mediation with LLM-Based Agents for Legal Research
- Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial
- Code2MCP: Transforming Code Repositories into MCP Services
- ProfilingAgent: Profiling-Guided Agentic Reasoning for Adaptive Model Optimization
- DualAlign: Generating Clinically Grounded Synthetic Data
- Internet 3.0: Architecture for a Web-of-Agents with it's Algorithm for Ranking Agents
- Emergent Social Dynamics of LLM Agents in the El Farol Bar Problem
- What Would an LLM Do? Evaluating Policymaking Capabilities of Large Language Models
- SasAgent: Multi-Agent AI System for Small-Angle Scattering Data Analysis
- Learning to Deliberate: Meta-policy Collaboration for Agentic LLMs with Multi-agent Reinforcement Learning
- Beyond Interpretability: Exploring the Comprehensibility of Adaptive Video Streaming through Large Language Models
- Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation
- The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices
- Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
- RumorSphere: A Framework for Million-scale Agent-based Dynamic Simulation of Rumor Propagation
- Batch Query Processing and Optimization for Agentic Workflows
- Graph RAG as Human Choice Model: Building a Data-Driven Mobility Agent with Preference Chain
- Plantbot: Integrating Plant and Robot through LLM Modular Agent Networks
- Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
- Heads or Tails: A Simple Example of Causal Abstractive Simulation
- From CVE Entries to Verifiable Exploits: An Automated Multi-Agent Framework for Reproducing CVEs
- LLM-Assisted Iterative Evolution with Swarm Intelligence Toward SuperBrain
- Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making
- The Resurgence of GCG Adversarial Attacks on Large Language Models
- Social World Models
- Synthetic Founders: AI-Generated Social Simulations for Startup Validation Research in Computational Social Science
- Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment
- COCORELI: Cooperative, Compositional Reconstitution & Execution of Language Instructions
- A Financial Brain Scan of the LLM
- ChatThero: An LLM-Supported Chatbot for Behavior Change and Therapeutic Support in Addiction Recovery
- Joint Enhancement of Relational Reasoning for Long-Context LLMs
- Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research
- Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
- Instructional Agents: LLM Agents on Automated Course Material Generation for Teaching Faculties
- Democracy-in-Silico: Institutional Design as Alignment in AI-Governed Polities
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- DeepMEL: A Multi-Agent Collaboration Framework for Multimodal Entity Linking
- EMNLP: Educator-role Moral and Normative Large Language Models Profiling
- From Bits to Boardrooms: A Cutting-Edge Multi-Agent LLM Framework for Business Excellence
- A Concurrent Modular Agent: Framework for Autonomous LLM Agents
- MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
- STARec: An Efficient Agent Framework for Recommender Systems via Autonomous Deliberate Reasoning
- Bias-Adjusted LLM Agents for Human-Like Decision-Making via Behavioral Economics
- Open-Universe Assistance Games
- Can LLM Agents Solve Collaborative Tasks? A Study on Urgency-Aware Planning and Coordination
- From Passive Tool to Socio-cognitive Teammate: A Conceptual Framework for Agentic AI in Human-AI Collaborative Learning
- Virtual Community: An Open World for Humans, Robots, and Society
- LLM-Powered Virtual Patient Agents for Interactive Clinical Skills Training with Automated Feedback
- CausalPlan: Empowering Efficient LLM Multi-Agent Collaboration Through Causality-Driven Planning
- `My Dataset of Love': A Preliminary Mixed-Method Exploration of Human-AI Romantic Relationships
- Personas within Parameters: Fine-Tuning Small Language Models with Low-Rank Adapters to Mimic User Behaviors
- HeroBench: A Benchmark for Long-Horizon Planning and Structured Reasoning in Virtual Worlds
- GTool: Graph Enhanced Tool Planning with Large Language Model
- Semantic Anchoring in Agentic Memory: Leveraging Linguistic Structures for Persistent Conversational Context
- Diagnostic-Guided Dynamic Profile Optimization for LLM-based User Simulators in Sequential Recommendation
- Do Large Language Model Agents Exhibit a Survival Instinct? An Empirical Study in a Sugarscape-Style Simulation
- Benchmarking LLM-based agents for single-cell omics analysis
- CHBench: A Cognitive Hierarchy Benchmark for Evaluating Strategic Reasoning Capability of LLMs
- Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework
- MM-Food-100K: A 100,000-Sample Multimodal Food Intelligence Dataset with Verifiable Provenance
- Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
- What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
- The PacifAIst Benchmark:Would an Artificial Intelligence Choose to Sacrifice Itself for Human Safety?
- Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
- The Roots of International Perceptions: Simulating US Attitude Changes Towards China with LLM Agents
- Simulating Generative Social Agents via Theory-Informed Workflow Design
- Exploring Large Language Model Agents for Piloting Social Experiments
- Cowpox: Towards the Immunity of VLM-based Multi-Agent Systems
- PersRM-R1: Enhance Personalized Reward Modeling with Reinforcement Learning
- IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
- Livia: An Emotion-Aware AR Companion Powered by Modular AI Agents and Progressive Memory Compression
- DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
- Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need
- Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation
- MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling
- FEAT: A Multi-Agent Forensic AI System with Domain-Adapted Large Language Model for Automated Cause-of-Death Analysis
- EMPATHIA: Multi-Faceted Human-AI Collaboration for Refugee Integration
- Agentic information systems
- Valid Inference with Imperfect Synthetic Data
- Scaling Personality Control in LLMs with Big Five Scaler Prompts
- LLMs for Resource Allocation: A Participatory Budgeting Approach to Inferring Preferences
- Simulating Human-Like Learning Dynamics with LLM-Empowered Agents
- The Term 'Agent' Has Been Diluted Beyond Utility and Requires Redefinition
- From MAS to MARS: Coordination Failures and Reasoning Trade-offs in Hierarchical Multi-Agent Robotic Systems within a Healthcare Scenario
- VirtLab: An AI-Powered System for Flexible, Customizable, and Large-scale Team Simulations
- DRAMA: A Dynamic and Robust Allocation-based Multi-Agent System for Changing Environments
- StackPilot: Autonomous Function Agents for Scalable and Environment-Free Code Execution
- Bridging Brains and Models: MoE-Based Functional Lesions for Simulating and Rehabilitating Aphasia
- ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"
- Galaxy: A Cognition-Centered Framework for Proactive, Privacy-Preserving, and Self-Evolving LLM Agents
- Sotopia-RL: Reward Design for Social Intelligence
- Putnam-AXIOM: A Functional and Static Benchmark for Measuring Higher Level Mathematical Reasoning in LLMs
- Parallelism Meets Adaptiveness: Scalable Documents Understanding in Multi-Agent LLM Systems
- Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
- Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
- Who is a Better Player: LLM against LLM
- AgentSME for Simulating Diverse Communication Modes in Smart Education
- Tree-of-Reasoning: Towards Complex Medical Diagnosis via Multi-Agent Reasoning with Evidence Tree
- Survey of Large Language Models in Extended Reality: Technical Paradigms and Application Frontiers
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
- InqEduAgent: Adaptive AI Learning Partners with Gaussian Process Augmentation
- Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired
- A Multi-Agent System for Complex Reasoning in Radiology Visual Question Answering
- Trainable Dynamic Mask Sparse Attention
- CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search
- Whispering Agents: An Event-driven Covert Communication Protocol For the Internet of Agents
- Meta-RAG on Large Codebases Using Code Summarization
- Social Media Information Operations
- L3M+P: Lifelong Planning with Large Language Models
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- Applying Psychometrics to Large Language Model Simulated Populations: Recreating the HEXACO Personality Inventory Experiment with Generative Agents
- From Logic to Language: A Trust Index for Problem Solving with LLMs
- A survey of multi-agent geosimulation methodologies: from ABM to LLM
- LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra
- DynaSwarm: Dynamically Graph Structure Selection for LLM-based Multi-agent System
- A Framework for Analyzing Abnormal Emergence in Service Ecosystems Through LLM-based Agent Intention Mining
- GasAgent: A Multi-Agent Framework for Automated Gas Optimization in Smart Contracts
- MetaAgent: Automatically Constructing Multi-Agent Systems Based on Finite State Machines
- Towards Simulating Social Influence Dynamics with LLM-based Multi-agents
- PersonaTwin: A Multi-Tier Prompt Conditioning Framework for Generating and Evaluating Personalized Digital Twins
- Mitigating Response Delays in Free-Form Conversations with LLM-powered Intelligent Virtual Agents
- An Explainable Emotion Alignment Framework for LLM-Empowered Agent in Metaverse Service Ecosystem
- Strategic Communication and Language Bias in Multi-Agent LLM Coordination
- FlowForge: Guiding the Creation of Multi-agent Workflows with Design Space Visualization as a Thinking Scaffold
- Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel Predictions
- MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
- A Multi-Agent Generative AI Framework for IC Module-Level Verification Automation
- Evaluation and Benchmarking of LLM Agents: A Survey
- A2HCoder: An LLM-Driven Coding Agent for Hierarchical Algorithm-to-HDL Translation
- MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations
- ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable Devices
- Games Agents Play: Towards Transactional Analysis in LLM-based Multi-Agent Systems
- Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
- On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- ChiMed 2.0: Advancing Chinese Medical Dataset in Facilitating Large Language Modeling
- MLC-Agent: Cognitive Model based on Memory-Learning Collaboration in LLM Empowered Agent Simulation Environment
- DynamiX: Large-Scale Dynamic Social Network Simulator
- Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
- CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback
- Integrating LLM in Agent-Based Social Simulation: Opportunities and Challenges
- Event-Driven Storytelling with Multiple Lifelike Humans in a 3D Scene
- PosterMate: Audience-driven Collaborative Persona Agents for Poster Design
- An Empirical Study of GenAI Adoption in Open-Source Game Development: Tools, Tasks, and Developer Challenges
- SYNTHIA: Synthetic Yet Naturally Tailored Human-Inspired PersonAs
- Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation
- A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
- Eywa: Provenance-Grounded Long-Term Memory for AI Agents
- Exploring the Dynamic Scheduling Space of Real-Time Generative AI Applications on Emerging Heterogeneous Systems
- A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents
- FutureSim: Replaying World Events to Evaluate Adaptive Agents
- SimPersona: Learning Discrete Buyer Personas from Raw Clickstreams for Grounded E-Commerce Agents
- Configurable multi-agent framework for scalable and realistic testing of llm-based agents
- Adaptive Multi-Agent Reasoning via Automated Workflow Generation
- Generative AI in Qualitative Research and Related Transparency Problems: A Novel Heuristic for Disclosing Uses of AI
- CogDual: Enhancing Dual Cognition of LLMs via Reinforcement Learning with Implicit Rule-Based Rewards
- Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learning
- Imitating Mistakes in a Learning Companion AI Agent for Online Peer Learning
- LLM Agent-Based Simulation of Student Activities and Mental Health Using Smartphone Sensing Data
- osmAG-LLM: Zero-Shot Open-Vocabulary Object Navigation via Semantic Maps and Large Language Models Reasoning
- Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
- Value-Based Large Language Model Agent Simulation for Mutual Evaluation of Trust and Interpersonal Closeness
- Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
- Can LLMs Generate Good Stories? Insights and Challenges from a Narrative Planning Perspective
- REVA: Supporting LLM-Generated Programming Feedback Validation at Scale Through User Attention-based Adaptation
- DiaryPlay: AI-Assisted Authoring of Interactive Vignettes for Everyday Storytelling
- Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
- Cultural Bias in Large Language Models: Evaluating AI Agents through Moral Questionnaires
- Large Population Models
- Infected Smallville: How Disease Threat Shapes Sociality in LLM Agents
- TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit
- Negotiating Comfort: Simulating Personality-Driven LLM Agents in Shared Residential Social Networks
- AgentsNet: Coordination and Collaborative Reasoning in Multi-Agent LLMs
- Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
- A Survey of Large Language Models in Discipline-specific Research: Challenges, Methods and Opportunities
- Hateful Person or Hateful Model? Investigating the Role of Personas in Hate Speech Detection by Large Language Models
- Multi-Actor Generative Artificial Intelligence as a Game Engine
- StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
- KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows
- State-Inference-Based Prompting for Natural Language Trading with Game NPCs
- MIND: A Multi-agent Framework for Zero-shot Harmful Meme Detection
- A Mathematical Theory of Discursive Networks
- Constella: Supporting Storywriters' Interconnected Character Creation through LLM-based Multi-Agents
- LLMs are Introvert
- MARBLE: A Multi-Agent Rule-Based LLM Reasoning Engine for Accident Severity Prediction
- FurniMAS: Language-Guided Furniture Decoration using Multi-Agent System
- ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
- Who's the Mole? Modeling and Detecting Intention-Hiding Malicious Agents in LLM-Based Multi-Agent Systems
- CREW-WILDFIRE: Benchmarking Agentic Multi-Agent Collaborations at Scale
- Zero-Mem: Zero-Token Memory Operations for LLM Agents
- The Agency Gap in AI-Supported Writing: How Reactive and Proactive Agent Designs Shape Multimodal Reasoning
- Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments
- Code Is the Body: Agent-Owned Software Bodies for Recursive Evolution and Descent
- Exploring a Gamified Personality Assessment Method through Interaction with LLM Agents Embodying Different Personalities
- Participatory Evolution of Artificial Life Systems via Semantic Feedback
- Leveraging Large Language Models for Tacit Knowledge Discovery in Organizational Contexts
- Towards Machine Theory of Mind with Large Language Model-Augmented Inverse Planning
- GRAFT: A Graph-based Flow-aware Agentic Framework for Document-level Machine Translation
- RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs
- The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models
- Can LLMs Play Ô Ăn Quan Game? A Study of Multi-Step Planning and Decision Making
- MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing
- Multi-Agent Reasoning for Cardiovascular Imaging Phenotype Analysis
- Synthetic Heuristic Evaluation: A Comparison between AI- and Human-Powered Usability Evaluation
- Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust
- AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness
- AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles
- Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications
- Generative Exaggeration in LLM Social Agents: Consistency, Bias, and Toxicity
- Ella: Embodied Social Agents with Lifelong Memory
- Towards the "Digital Me": A vision of authentic Conversational Agents powered by personal Human Digital Twins
- Epitome: Pioneering an Experimental Platform for AI-Social Science Integration
- Cognitive Weave: Synthesizing Abstracted Knowledge with a Spatio-Temporal Resonance Graph
- Hierarchical Memory Organization for Wikipedia Generation
- GATSim: Urban Mobility Simulation with Generative Agents
- Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games
- Knowledge Augmented Finetuning Matters in both RAG and Agent Based Dialog Systems
- Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism
- Universal Retrieval for Multimodal Trajectory Modeling
- GenEscape: Hierarchical Multi-Agent Generation of Escape Room Puzzles
- MobiVerse: Scaling Urban Mobility Simulation with Hybrid Lightweight Domain-Specific Generator and Large Language Models
- Gradient-Based Neuroplastic Adaptation for Concurrent Optimization of Neuro-Fuzzy Networks
- G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
- A Literature Review on Simulation in Conversational Recommender Systems
- Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
- LLM-Based Social Simulations Require a Boundary
- Augmenting Multi-Agent Communication with State Delta Trajectory
- TRIZ Agents: A Multi-Agent LLM Approach for TRIZ-Based Innovation
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
- Breaking Single-Tester Limits: Multi-Agent LLMs for Multi-User Feature Testing
- MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
- Large Language Models as Psychological Simulators: A Methodological Guide
- From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology
- PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases
- Arch-Router: Aligning LLM Routing with Human Preferences
- SimuPanel: A Novel Immersive Multi-Agent System to Simulate Interactive Expert Panel Discussion
- AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need
- Doppelganger Method: Breaking Role Consistency in LLM Agent via Prompt-based Transferable Adversarial Attack
- SimSpark: Interactive Simulation of Social Media Behaviors
- StorySage: Conversational Autobiography Writing Powered by a Multi-Agent Framework
- Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
- Modeling Earth-Scale Human-Like Societies with One Billion Agents
- From What to Respond to When to Respond: Timely Response Generation for Open-domain Dialogue Agents
- LocationReasoner: Evaluating LLMs on Real-World Site Selection Reasoning
- Confident-Knowledge Diversity Drives Human-Human and Human-AI Free Discussion Synergy and Reveals Pure-AI Discussion Shortfalls
- Evaluating Cell Type Inference in Vision Language Models Under Varying Visual Context
- Behavioral Generative Agents for Energy Operations
- Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics
- AgentSense: Virtual Sensor Data Generation Using LLM Agents in Simulated Home Environments
- EconGym: A Scalable AI Testbed with Diverse Economic Tasks
- Black-Box Access is Insufficient for Rigorous AI Audits
- From Emergence to Control: Probing and Modulating Self-Reflection in Language Models
- PE-MA: Parameter-Efficient Co-Evolution of Multi-Agent Systems
- A Study on Individual Spatiotemporal Activity Generation Method Using MCP-Enhanced Chain-of-Thought Large Language Models
- Can LLMs Reason About Trust?: A Pilot Study
- AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
- MAPLE: Multi-Agent Adaptive Planning with Long-Term Memory for Table Reasoning
- MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks
- Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
- OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
- Teaming in the AI Era: AI-Augmented Frameworks for Forming, Simulating, and Optimizing Human Teams
- Evaluating Prompt-Driven Chinese Large Language Models: The Influence of Persona Assignment on Stereotypes and Safeguards
- Gen-n-Val: Agentic Image Data Generation and Validation
- Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling
- Truly Self-Improving Agents Require Intrinsic Metacognitive Learning
- Time to Talk: LLM Agents for Asynchronous Group Communication in Mafia Games
- SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents
- Using Large Language Models to Simulate Human Behavioural Experiments: Port of Mars
- Towards Language-Augmented Multi-Agent Deep Reinforcement Learning
- Benchmarking LLMs' Swarm intelligence
- Optimization Problem Solving Can Transition to Evolutionary Agentic Workflows
- AI Agent Behavioral Science
- VChatter: Exploring Generative Conversational Agents for Simulating Exposure Therapy to Reduce Social Anxiety
- Simulating Human Behavior with the Psychological-mechanism Agent: Integrating Feeling, Thought, and Action
- AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
- Decompose, Plan in Parallel, and Merge: A Novel Paradigm for Large Language Models based Planning with Multiple Constraints
- Mapping Student-AI Interaction Dynamics in Multi-Agent Learning Environments: Supporting Personalised Learning and Reducing Performance Gaps
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
- MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching
- Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
- An Empirical Study of Group Conformity in Multi-Agent Systems
- Comprehensive Vulnerability Analysis is Necessary for Trustworthy LLM-MAS
- Will artificial agents pursue power by default?
- Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
- Not All Jokes Land: Evaluating Large Language Models Understanding of Workplace Humor
- Beyond Static Responses: Multi-Agent LLM Systems as a New Paradigm for Social Science Research
- HASHIRU: Hierarchical Agent System for Hybrid Intelligent Resource Utilization
- Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
- Artificial Behavior Intelligence: Technology, Challenges, and Future Directions
- MIRROR: Converging Cognitive Principles as Computational Mechanisms for AI Reasoning
- Reasoning Like an Economist: Post-Training on Economic Problems Induces Strategic Generalization in LLMs
- DefenderBench: A Toolkit for Evaluating Language Agents in Cybersecurity Environments
- When Harry Meets Superman: The Role of The Interlocutor in Persona-Based Dialogue Generation
- Exploring the Impact of Occupational Personas on Domain-Specific QA
- The Power of Stories: Narrative Priming Shapes How LLM Agents Collaborate and Compete
- Finance as Extended Biology: Reciprocity as the Cognitive Substrate of Financial Behavior
- Redefining Research Crowdsourcing: Incorporating Human Feedback with LLM-Powered Digital Twins
- Cross-Task Experiential Learning on LLM-based Multi-Agent Collaboration
- When artificial intelligence substitutes humans in higher education: the cost of loneliness, student success, and retention
- PhotoArtAgent: Intelligent Photo Retouching with Language Model-Based Artist Agents
- LLM Agents for Bargaining with Utility-based Feedback
- Large Language Model-Based Agents for Automated Research Reproducibility: An Exploratory Study in Alzheimer's Disease
- The Cognitive Foundations of Economic Exchange: A Modular Framework Grounded in Behavioral Evidence
- Be.FM: Open Foundation Models for Human Behavior
- Free Lunch for User Experience: Crowdsourcing Agents for Scalable User Studies
- Sentiment Simulation using Generative AI Agents
- Retweets, Receipts, and Resistance: Discourse, Sentiment, and Credibility in Public Health Crisis Twitter
- ValueSim: Generating Backstories to Model Individual Value Systems
- Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development
- Beyond Monoliths: Expert Orchestration for More Capable, Democratic, and Safe Language Models
- Risks of AI-driven product development and strategies for their mitigation
- Herd Behavior: Investigating Peer Influence in LLM-based Multi-Agent Systems
- CoderAgent: Simulating Student Behavior for Personalized Programming Learning with Large Language Models
- Public Discourse Sandbox: Facilitating Human and AI Digital Communication Research
- Large Language Models Miss the Multi-Agent Mark
- GGBond: Growing Graph-Based AI-Agent Society for Socially-Aware Recommender Simulation
- MultiPhishGuard: An LLM-based Multi-Agent System for Phishing Email Detection
- CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic
- AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
- Recalibrating the Compass: Integrating Large Language Models into Classical Research Methods
- From Single to Multi-Granularity: Toward Long-Term Memory Association and Selection of Conversational Agents
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- Multi-Agent Collaboration via Evolving Orchestration
- Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
- MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
- OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
- Generating HomeAssistant Automations Using an LLM-based Chatbot
- Agents Require Metacognitive and Strategic Reasoning to Succeed in the Coming Labor Markets
- CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems
- Agentic Visualization: Extracting Agent-based Design Patterns from Visualization Systems
- Position: Collaborative Agentic AI Needs Interoperability Across Ecosystems
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems
- Aligning LLM with human travel choices: a persona-based embedding learning approach
- When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary
- Response Uncertainty and Probe Modeling: Two Sides of the Same Coin in LLM Interpretability?
- RoleRAG: Enhancing LLM Role-Playing via Graph Guided Retrieval
- Survival Games: Human-LLM Strategic Showdowns under Severe Resource Scarcity
- FedWorld: Scope-Aware Federation of Agent World Models
- The Real Barrier to LLM Agent Usability is Agentic ROI
- Self-Improving Large Language Models via Progressive Experience Evolution
- Understanding How Value Neurons Shape the Generation of Specified Values in LLMs
- Rethinking Agent Design: From Top-Down Workflows to Bottom-Up Skill Evolution
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
- Runaway is Ashamed, But Helpful: On the Early-Exit Behavior of Large Language Model-based Agents in Embodied Environments
- Large language model as user daily behavior data generator: balancing population diversity and individual personality
- Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions
- MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents
- RetroChat: Designing for the Preservation of Past Digital Experiences
- KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model Improvement
- CoEvo-Mem: Co-Evolving Retrieval Policy and Memory Bank for LLM Agents
- Advancing the Scientific Method with Large Language Models: From Hypothesis to Discovery
- Optimizing LLM-Based Multi-Agent System with Textual Feedback: A Case Study on Software Development
- MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents
- MASLab: A Unified and Comprehensive Codebase for LLM-based Multi-Agent Systems
- How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior
- Salami Attack: Stealthy Collusive Memory Poisoning against OpenClaw
- No One Wins in Nuclear War: A Social Simulation of Military Decision-making
- Evolving in the Agent Jungle via History-Informed Opponent Awareness
- Position: Agentic Systems Constitute a Key Component of Next-Generation Intelligent Image Processing
- AGENT-X: Adaptive Guideline-based Expert Network for Threshold-free AI-generated teXt detection
- P2VA: Converting Persona Descriptions into Voice Attributes for Fair and Controllable Text-to-Speech
- AutoData: A Multi-Agent System for Open Web Data Collection
- How Managers Perceive AI-Assisted Conversational Training for Workplace Communication
- Concept Incongruence: An Exploration of Time and Death in Role Playing
- Modeling Social Dynamics with an LLM-Enabled Agent Based Network-Dynamic (LAND) Model
- Simulation Agent: A Framework for Integrating Simulation and Large Language Models for Enhanced Decision-Making
- MAFA: A multi-agent framework for annotation
- PsyMem: Fine-grained psychological alignment and Explicit Memory Control for Advanced Role-Playing LLMs
- Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
- G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning
- MRAFnd: Multimodal Retrieval-Augmented Framework for Zero-Shot Fake News Detection
- CAIM: Development and Evaluation of a Cognitive AI Memory Framework for Long-Term Interaction with Intelligent Agents
- InnateCoder: Learning Programmatic Options with Foundation Models
- Automated Profile Inference with Language Model Agents
- ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors
- Interactional Fairness in LLM Multi-Agent Systems: An Evaluation Framework
- PMMC: Prospective Multimodal Memory Compilation for Long-Term LVLM Agents
- TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents
- LLM Agents Are Hypersensitive to Nudges
- XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration
- MAPLE-Guard: Memory-Aware Link Enforcement Against Memory-Link Poisoning in Multi-Agent Systems
- Creating General User Models from Computer Use
- Artificial Intelligence and Modeling & Simulation: An Overview
- Systematic Failures in Collective Reasoning under Distributed Information in Multi-Agent LLMs
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
- S4R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching
- Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments
- Characterizing Unintended Consequences in Human-GUI Agent Collaboration for Web Browsing
- Design and Evaluation of Generative Agent-based Platform for Human-Assistant Interaction Research: A Tale of 10 User Studies
- Interpretable Risk Mitigation in LLM Agent Systems
- From We to Me: Theory Informed Narrative Shift with Abductive Reasoning
- MASS: Muli-agent simulation scaling for portfolio construction
- Trustless Autonomy: Understanding Motivations, Benefits, and Governance Dilemmas in Self-Sovereign Decentralized AI Agents
- SALM: A Multi-Agent Framework for Language Model-Driven Social Network Simulation
- LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations
- The Influence of Human-inspired Agentic Sophistication in LLM-driven Strategic Reasoners
- The Truth Becomes Clearer Through Debate! Multi-Agent Systems with Large Language Models Unmask Fake News
- Agent-as-a-Service based on Agent Network
- Reciprocity as the Foundational Substrate of Society: How Reciprocal Dynamics Scale into Social Systems
- CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents
- Memory Reward Inflation in Self-Improving LLM Agents
- AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategies in Conversational AI Agents
- MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents
- Putting It All into Context: Simplifying Agents with LCLMs
- Internet of Agents: Fundamentals, Applications, and Challenges
- PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications
- Artificial intelligence and free will: generative agents utilizing large language models have functional free will
- Can Generative AI agents behave like humans? Evidence from laboratory market experiments
- EcoLANG: Efficient and Effective Agent Communication Language Induction for Social Simulation
- Applying Cognitive Design Patterns to General LLM Agents
- Reputation as a Solution to Cooperation Collapse in LLM-based MASs
- Do MLLMs Capture How Interfaces Guide User Behavior? A Benchmark for Multimodal UI/UX Design Understanding
- Exploring Silicon-Based Societies: An Early Study of the Moltbook Agent Community
- MemEngine: A Unified and Modular Library for Developing Advanced Memory of LLM-based Agents
YC-Bench: Benchmarking AI Agents for Long-Term Planning and Consistent Execution- Co3Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion
- AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents
- Artificial Intelligence in Government: Why People Feel They Lose Control
- Towards Autonomous Micromobility through Scalable Urban Simulation
- The Anatomy of the Moltbook Social Graph
- Self-Evolving Multi-Agent Framework for Efficient Decision Making in Real-Time Strategy Scenarios
- Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
- Habermolt: Delegating Deliberation to AI Representatives
- Social Theory Should Be a Structural Prior for Agentic AI: A Formal Framework for Multi-Agent Social Systems
- Vibe Researching as Wolf Coming: Can AI Agents with Skills Replace or Augment Social Scientists?
- Holos: A Web-Scale LLM-Based Multi-Agent System for the Agentic Web
- Behavioral Indicators of Overreliance During Interaction with Conversational Language Models
- When Workout Buddies Are Virtual: AI Agents and Human Peers in a Longitudinal Physical Activity Study
- S3: Improving Agent Safety through Multi-Stage Defense
- Everyone Conforms, No One Believes: Pluralistic Ignorance in LLM Agent Populations
- Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, Cross-Principal Approach
- Quo Vadis, World Modeling?
- Governable Individuals: An Identity Layer for Embodied Agents That Keep Learning
- Attractor States Emerge in Multi-Turn LLM Conversations
- D-MEM: Dopamine-Gated Agentic Memory via Reward Prediction Error Routing
- Collective Cognition in Hybrid Groups: A Network Science Synthesis
- AI Agents in Financial Markets: Architecture, Applications, and Systemic Implications
- STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Stories
- "Humans welcome to observe": A First Look at the Agent Social Network Moltbook
- Do Language Models Pass the Bechdel Test? Auditing Gender Biases in LLM-Generated Screenplays
- Large Language Models Do Not Always Need Readable Language
- Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation
- Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation
- MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
- AgentSpec: Understanding Embodied Agent Scaffolds Through Controlled Composition
- EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents
- GRPO Does Not Close the Multi-Agent Coordination Gap
- Toward Epistemic Stability: Engineering Consistent Procedures for Industrial LLM Hallucination Reduction
- Memory for Autonomous LLM Agents:Mechanisms, Evaluation, and Emerging Frontiers
- A Miniature Brain Transformer: Thalamic Gating, Hippocampal Lateralization, Amygdaloid Salience, and Prefrontal Working Memory in Attention-Coupled Latent Memory
- Real-Time AI Service Economy: A Framework for Agentic Computing Across the Continuum
- Theory Discovery in Social Networks: Automating ERGM Specification with Large Language Models
- Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation
- Recursive Models for Long-Horizon Reasoning
- Persistent Identity in AI Agents: A Multi-Anchor Architecture for Resilient Memory and Continuity
- Data Therapist: Eliciting Domain Knowledge from Subject Matter Experts Using Large Language Models
- RAVEL: Reasoning Agents for Validating and Evaluating LLM Text Synthesis
- Toward Expert Investment Teams:A Multi-Agent LLM System with Fine-Grained Trading Tasks
- ParamMem: Augmenting Language Agents with Parametric Reflective Memory
- Positive Alignment: Artificial Intelligence for Human Flourishing
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- SkillNet: Create, Evaluate, and Connect AI Skills
- FactorMiner: A Self-Evolving Agent with Skills and Experience Memory for Financial Alpha Discovery
- When AI Agents Teach Each Other: Discourse Patterns Resembling Peer Learning in the Moltbook Community
- Language model agents show in-group trust bias invisible to standard behavioural audits
- Causal methods for LLM development and evaluation
- The Model Is Not the Product: A Dual-Pillar Architecture for Local-First Psychological Coaching
- MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structural Anomaly Detection
- Fast Response or Silence: Conversation Persistence in an AI-Agent Social Network
- A New Strategy for Artificial Intelligence: Training Foundation Models Directly on Human Brain Data
- On the Failure of Latent State Persistence in Large Language Models
- Can Memory-Augmented LLM Agents Aid Journalism in Interpreting and Framing News for Diverse Audiences?
- Characterizing AI Agents for Alignment and Governance
- MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model Framework
- MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
- S-Bus: Automatic Read-Set Reconstruction for Multi-Agent LLM State Coordination
- Turning Intent into Specifications: A Benchmark and an Interactive User-Assistant Agent
- Agentic Test-Time Scaling for WebAgents
- Engineering-Oriented Symbolic Regression: LLMs as Physics Agents for Discovery of Simulation-Ready Constitutive Laws
- The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents
- Adaptivity Under Realizability Constraints: Comparing In-Context and Agentic Learning
- MRMS: A Multi-Resolution Memory Substrate for Long-Lived AI Agents
- Search-Based Interaction For Conversation Recommendation via Generative Reward Model Based Simulated User
- Security Threat Modeling for Emerging AI-Agent Protocols: A Comparative Analysis of MCP, A2A, Agora, and ANP
- From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review
- When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems
- Contagion Networks: Evaluator Preference Propagation in Multi-Agent LLM Systems
- Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems
- Deployment-Time Memorization in Foundation-Model Agents
- Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads
- From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents
- ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactions
- Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance
- Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents
- AgenticAITA: A Proof-Of-Concept About Deliberative Multi-Agent Reasoning for Autonomous Trading Systems
- GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)
- Beneath the Surface: Investigating LLMs' Capabilities for Communicating with Subtext
- Omni-SimpleMem: Autoresearch-Guided Discovery of Lifelong Multimodal Agent Memory
- Interpretable Context Methodology: Folder Structure as Agentic Architecture
- Talk, Judge, Cooperate: Gossip-Driven Indirect Reciprocity in Self-Interested LLM Agents
- HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense
- Evolution of Cooperation in LLM-Agent Societies: A Preliminary Study Using Different Punishment Strategies
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
- OpenFOAMGPT 2.0: end-to-end, trustworthy automation for computational fluid dynamics
- Mesh Memory Protocol: Semantic Infrastructure for Multi-Agent LLM Systems
- HorizonBench: Long-Horizon Personalization with Evolving Preferences
- Dynamics of Cognitive Heterogeneity: Investigating Behavioral Biases in Multi-Stage Supply Chains with LLM-Based Simulation
- Situation Graph Prediction: Structured Perspective Inference for User Modeling
- The effects of generative AI agents and scaffolding on enhancing students’ comprehension of visual learning analytics
- Generative AI Literacy: A Comprehensive Framework for Literacy and Responsible Use
- Exploring Personality-Aware Interactions in Salesperson Dialogue Agents
- The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
- AI-based Verbal and Visual Scaffolding in a Serious Game: Effects on Learning and Cognitive Load
- TeamLLM: Exploring the Capabilities of LLMs for Multimodal Group Interaction Prediction
- Designing Digital Humans with Ambient Intelligence
- SoK: Blockchain Agent-to-Agent Payments
- Debiasing LLMs by Fine-tuning
- GTA: Generative Traffic Agents for Simulating Realistic Mobility Behavior
- Exploring Implicit Perspectives on Autism in Large Language Models Through Multi-Agent Simulations
- AI Agents Need Memory Control Over More Context
- Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents
- DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents
- AI Agent Economics: Can Autonomous Economic Behavior Emerge among AI Agents under Minimal External Conditions?
- Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation
- When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
- ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning
- WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
- From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities
- LoongFlow: Directed Evolutionary Search via a Cognitive Plan-Execute-Summarize Paradigm
- MAGI: Multi-Agent Guided Interview for Psychiatric Assessment
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- AI Awareness
- Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning
- Preference-Driven Online Adaptation for Personalized Interaction Initiation in Proactive AI Assistants
- Blockchain Empowered Trustworthy Agent Networks: Foundations, Taxonomy, and Future Directions
- Caching for the Future: Scrub Jay Episodic Memory Principles for Agent Memory Systems
- A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination
- MatrAIx: Simulating the World with 8.3 Billion Persona Agents
- FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents
- Contextual Agentic Memory is a Memo, Not True Memory
- Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems
- LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment
- Artificial Institutions: How Institutional Design Shapes LLM Simulations
- CogniFold: Always-On Proactive Memory via Cognitive Folding
- XGrammar-2: Dynamic and Efficient Structured Generation Engine for Agentic LLMs
- TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
- Cognitive Silicon: An Architectural Blueprint for Post-Industrial Computing Systems
- PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression
- IMPersona: Evaluating Individual Level LM Impersonation
- Building LLM Agents by Incorporating Insights from Computer Systems
- From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs
- Reflexive Prompt Engineering: A Framework for Responsible Prompt Engineering and Interaction Design
- Interpretable Locomotion Prediction in Construction Using a Memory-Driven LLM Agent With Chain-of-Thought Reasoning
- FlowReasoner: Reinforcing Query-Level Meta-Agents
- EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue Framework
- Exploring Collaborative GenAI Agents in Synchronous Group Settings: Eliciting Team Perceptions and Design Considerations for the Future of Work
- NoWag: A Unified Framework for Shape Preserving Compression of Large Language Models
- BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation
- AI with Emotions: Exploring Emotional Expressions in Large Language Models
- Understanding the Repeat Curse in Large Language Models from a Feature Perspective
- FAIRGAME: a Framework for AI Agents Bias Recognition using Game Theory
- SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulation
- DashChat: Interactive Authoring of Performance Dashboard Design Prototypes through Conversation with LLM-Powered Agent
- The Athenian Academy: A Seven-Layer Architecture Model for Multi-Agent Systems
- From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems
- Mind the Gaps: Mixture-of-Minds for Human Simulation
- When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents
- Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
- ASIDE: From Conflict Participants to Co-Observers Through Dyadic Spectator Reflection
- Causal Episodic Memory for Feedback-Driven Agent Repair
- APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning
- Persona-Pruner: Sculpting Lightweight Models for Role-Playing
- DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory Graph
- Memory in the LLM Era: Modular Architectures and Strategies in a Unified Framework
- Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data
- Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection
- Adaptive Human-Agent Teaming: A Review of Empirical Studies from the Process Dynamics Perspective
- Learning to Be A Doctor: Searching for Effective Medical Agent Architectures
- LLM-Driven NPCs: Cross-Platform Dialogue System for Games and Social Platforms
- SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World Users
- A Survey of Personalization: From RAG to Agent
- C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation
- AgentDynEx: Nudging the Mechanics and Dynamics of Multi-Agent Simulations
- EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety
- UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
- Semantic Commit: Helping Users Update Intent Specifications for AI Memory at Scale
- Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare
- Deceptive Automated Interpretability: Language Models Coordinating to Fool Oversight Systems
- MOSAIC: Modeling Social AI for Content Dissemination and Regulation in Multi-Agent Simulations
- Exploring Human-Like Thinking in Search Simulations with Large Language Models
- TALE: A Tool-Augmented Framework for Reference-Free Evaluation of Large Language Models
- Toward Holistic Evaluation of Recommender Systems Powered by Generative Models
- Review of Case-Based Reasoning for LLM Agents: Theoretical Foundations, Architectural Components, and Cognitive Integration
- V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
- Agent Guide: A Simple Agent Behavioral Watermarking Framework
- Can LLMs Simulate Personas with Reversed Performance? A Systematic Investigation for Counterfactual Instruction Following in Math Reasoning Context
- A Desideratum for Conversational Agents: Capabilities, Challenges, and Future Directions
- Stock Market Forecasting: From Traditional Predictive Models to Large Language Models
- DoCIA: An Online Document-Level Context Incorporation Agent for Speech Translation
- EduPlanner: LLM-Based Multi-Agent Systems for Customized and Intelligent Instructional Design
- Agent-based model [wikipedia]
Discussions
- Generative Agents: Interactive Simulacra of Human Behavior [hn, 391 points, 252 comments]
- 1. A large and growing literature on "generative agent-based modeling" aims to study social science questions by using LLMs to simulate the behavior of people in naturalistic situations. I think it's [bsky, 193 points, 12 comments]
- Generative Agents: Interactive Simulacra of Human Behavior [hn, 13 points, 2 comments]
- New Advanced Video Game AI - Generative Agents: Interactive Simulacra of Human Behavior [lemmy, 13 points, 1 comments]
- Like, people have been working on this kind of problem for decades. Chris Crawford even galloped out of GDC on an imaginary horse and shunned the entire rest of the game industry for decades to work o [bsky, 7 points, 1 comments]
- אין לי כח לכל הדיבור על moltbook הרשת החברתית של הבוטים הנוראיים (בניגוד ל-X, הרשת החברתית של הבוטים הנוראיים). אבל לרקע - המאמר הזה מלפני שנתיים וחצי: arxiv.org/abs/2304.03442 [bsky, 6 points, 1 comments]
- My position is that agentic AI relies on “sharp edged” problems for which there are big costs if it makes a mistake, implying very high accuracy requirements—and for most agentic uses, we aren’t at th [bsky, 3 points, 1 comments]
- I have a few models in the works, but this one is inspired by this "Sims" proof of concept where LLM-powered agents interact quite freely with each other. Really interesting paper. As I said, just wat [bsky, 3 points, 1 comments]
- You can read the paper here: arxiv.org/pdf/2304.03442 [bsky, 3 points, 1 comments]
- [R] Generative Agents: Interactive Simulacra of Human Behavior (to be presented at UIST) [lemmy, 3 points, 0 comments]
- Generative Agents: Interactive Simulacra of Human Behavior [lobsters, 2 points, 0 comments]
- Oh ye, I remember I had to look into that for work. I think you are talking about this: arxiv.org/pdf/2304.03442? The outcome was pretty impressive, but there were some severe limitations in time of [bsky, 2 points, 1 comments]
- arxiv.org/abs/2304.03442 [bsky, 1 points, 0 comments]
- In case you missed it, Stanford published a paper about AI agents: https://arxiv.org/pdf/2304.03442.pdf Generative agents wake up, cook breakfast, head to work; They form opinions, notice each other, [bsky, 1 points, 0 comments]
- Fascinating stuff. I’m still reading through your article so perhaps you already mention this, but have you seen this paper? arxiv.org/pdf/2304.03442 [bsky, 1 points, 2 comments]
- 🎧 EP058: Inside the Autonomous AI Town of Smallville 📄 Generative Agents 🔗 https://arxiv.org/abs/2304.03442 🟢 https://podcasters.spotify.com/pod/show/yun-wu/episodes/EP058-Inside-the-Autonomous-AI [bsky, 1 points, 0 comments]
- Stanford and Google researchers released 25 AI bots into a virtual town, Smallville. These agents cooked, worked, and socialized like humans. They roamed schools, cafés, and bars - like The Sims but w [bsky, 1 points, 0 comments]
- This paper has a good and simple memory consolidation mechanism: arxiv.org/abs/2304.03442 [bsky, 1 points, 1 comments]
- This is the paper btw. arxiv.org/abs/2304.03442 I remember it because it was part of a seminar and looked interesting enough for the not so LLM skeptic me at first glance. How things have changed in t [bsky, 1 points, 0 comments]
- Check out this research paper from three years ago then. Even the earlier models were already pretty good at simulating basic human behavior with the right setup. arxiv.org/pdf/2304.03442 [bsky, 0 points, 0 comments]
- Generative Agents: Interactive Simulacra of Human Behavior arxiv.org/abs/2304.03442 [bsky, 0 points, 0 comments]
- I heard something similar a while ago,but it was not on Minecraft arxiv.org/abs/2304.034... It is really interesting I wonder if we would ever get NPC towns with this level of involvement [bsky, 0 points, 0 comments]
- At Stanford, a research team has created a city simulation wherein ChatGPT-trained “generative agents” appear to approximate “plausible” human behavior. Notably, they had “meaningful” conversations [bsky, 0 points, 1 comments]
- This is one of my favorite AI papers recently uses a network of LLMs to simulate a Stardew valley esque community. https://arxiv.org/abs/2304.03442 They observed some pretty incredible emergent soci [bsky, 0 points, 0 comments]
- Westworld vibes 🦄 https://arxiv.org/pdf/2304.03442.pdf [bsky, 0 points, 0 comments]
- @theophite.bsky.social do you know of any actual research backing this stuff or is it just utter nonsense? Lile they are partnering with this company Simile, and their only publication on their site i [bsky, 0 points, 0 comments]
Related