Designing Digital Humans with Ambient Intelligence
2026/04/06 by Mengyu Chen, Pranav Deshpande, Runqing Yang +4 · 1 voice
Computer Science · #cs.HC #cs.MA
paper · pdf · doi:10.48550/arxiv.2604.05120
arxiv published 2026/04/06 · arxiv updated 2026/04/29
Abstract
Digital humans are lifelike virtual agents capable of natural conversation and are increasingly deployed in domains like retail and finance. However, most current digital humans operate in isolation from their surroundings and lack contextual awareness beyond the dialogue itself. We address this limitation by integrating ambient intelligence (AmI) - i.e., environmental sensors, IoT data, and contextual modeling - with digital human systems. This integration enables situational awareness of the user's environment, anticipatory and proactive assistance, seamless cross-device interactions, and personalized long-term user support. We present a conceptual framework defining key roles that AmI can play in shaping digital human behavior, a design space highlighting dimensions such as proactivity levels and privacy strategies, and application-driven patterns with case studies in financial and retail services. We also discuss an architecture for ambient-enabled digital humans and provide guidelines for responsible design regarding privacy and data governance. Together, our work positions ambient intelligent digital humans as a new class of interactive agents powered by AI that respond not only to users' queries but also to the context and situations in which the interaction occurs.
Citations
- A Survey of Context Engineering for Large Language Models
- Proactive Conversational AI: A Comprehensive Survey of Advancements and Opportunities
- WorldSimBench: Towards Video Generation Models as World Simulators
- Enabling Data-Driven and Empathetic Interactions: A Context-Aware 3D Virtual Agent in Mixed Reality for Enhanced Financial Customer Experience
- The Llama 3 Herd of Models
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- OpenVLA: An Open-Source Vision-Language-Action Model
- A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems
- Evaluating Very Long-Term Conversational Memory of LLM Agents
- Genie: Generative Interactive Environments
- Multi-View Conformal Learning for Heterogeneous Sensor Fusion
- Learning Interactive Real-World Simulators
- The Rise and Potential of Large Language Model Based Agents: A Survey
- Benchmarking Large Language Models in Retrieval-Augmented Generation
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- 3D Gaussian Splatting for Real-Time Radiance Field Rendering
- Autonomous Tester Agent Benchmark
- Mind2Web: Towards a Generalist Agent for the Web
- Gorilla: Large Language Model Connected with Massive APIs
- Generative Agents: Interactive Simulacra of Human Behavior
- GPT-4 Technical Report
- PaLM-E: An Embodied Multimodal Language Model
- End-to-End Speech Recognition: A Survey
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Accountability in artificial intelligence: what it is and how it works
- Robust Speech Recognition via Large-Scale Weak Supervision
- Distributing Accountability, Not Capability: Phase Separation and the LLM Workflow Quadrant in Autonomous AI Agent Architectures
- Transformers are Sample-Efficient World Models
- Training language models to follow instructions with human feedback
- BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
- ByteTrack: Multi-Object Tracking by Associating Every Detection Box
- ByteTrack: Multi-object Tracking by Associating Every Detection Box
- AI, big data, and the future of consent
- Attention Bottlenecks for Multimodal Fusion
- AdaMML: Adaptive Multi-Modal Learning for Efficient Video Recognition
- Recent Advances in Deep Learning Based Dialogue Systems: A Systematic Survey
- Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar Reconstruction
- Denoising Diffusion Probabilistic Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Neural Voice Puppetry: Audio-driven Facial Reenactment
- Advances and Open Problems in Federated Learning
- Dream to Control: Learning Behaviors by Latent Imagination
- pyannote.audio: neural building blocks for speaker diarization
- Deep Evidential Regression
- Artificial Intelligence: the global landscape of ethics guidelines
- MediaPipe: A Framework for Building Perception Pipelines
- Capture, Learning, and Synthesis of 3D Speaking Styles
- A Style-Based Generator Architecture for Generative Adversarial Networks
- A Style-Based Generator Architecture for Generative Adversarial Networks
- Deep Appearance Models for Face Rendering
- Evidential Deep Learning to Quantify Classification Uncertainty
- World Models
- A Survey of Indoor Localization Systems and Technologies
- A review of empirical evidence on different uncanny valley hypotheses: support for perceptual mismatch as one road to the valley of eeriness
- It’s only a computer: Virtual humans increase willingness to disclose
- Context Aware Computing for The Internet of Things: A Survey
- Internet of Things (IoT): A Vision, Architectural Elements, and Future Directions
- Internet of Things (IoT): A vision, architectural elements, and future directions
- Place illusion and plausibility can lead to realistic behaviour in immersive virtual environments
- The Effect of the Agency and Anthropomorphism on Users' Sense of Telepresence, Copresence, and Social Presence in Virtual Environments
- Reducing consistency in human realism increases the uncanny valley effect; increasing category uncertainty does not
- A tutorial on human activity recognition using body-worn inertial sensors
Discussions
Related