ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems
2025/12/07 by Yang, Bufang, Xu, Lilin, Zeng, Liekang +8
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC)
paper · doi:10.48550/arxiv.2512.06721
Abstract
Large Language Model (LLM) agents are emerging to transform daily life. However, existing LLM agents primarily follow a reactive paradigm, relying on explicit user instructions to initiate services, which increases both physical and cognitive workload. In this paper, we propose ProAgent, the first end-to-end proactive agent system that harnesses massive sensory contexts and LLM reasoning to deliver proactive assistance. ProAgent first employs a proactive-oriented context extraction approach with on-demand tiered perception to continuously sense the environment and derive hierarchical contexts that incorporate both sensory and persona cues. ProAgent then adopts a context-aware proactive reasoner to map these contexts to user needs and tool calls, providing proactive assistance. We implement ProAgent on Augmented Reality (AR) glasses with an edge server and extensively evaluate it on a real-world testbed, a public dataset, and through a user study. Results show that ProAgent achieves up to 33.4% higher proactive prediction accuracy, 16.8% higher tool-calling F1 score, and notable improvements in user satisfaction over state-of-the-art baselines, marking a significant step toward proactive assistants. A video demonstration of ProAgent is available at https://youtu.be/pRXZuzvrcVs.
Citations
- Sensible Agent: A Framework for Unobtrusive Interaction with Proactive AR Agents
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
- EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses
- ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable Devices
- DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
- FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
- ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- CodingGenie: A Proactive LLM-Powered Programming Assistant
- AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications
- WatchGuardian: Enabling User-Defined Personalized Just-in-Time Intervention on Smartwatch
- SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions
- EmbedGenius: Towards Automated Software Development for Generic Embedded IoT Systems
- SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live Interactions
- A Survey on LLM-as-a-Judge
- Satori: Towards Proactive AR Assistant with Belief-Desire-Intention User Modeling
- Proactive Agent: Shifting LLM Agents from Reactive Responses to Active Assistance
- Scaling Synthetic Data Creation with 1,000,000,000 Personas
- Towards Open Respiratory Acoustic Foundation Models: Pretraining and Benchmarking
- Ask-before-Plan: Proactive Language Agents for Real-World Planning
- Aligning LLM Agents by Learning Latent Preference from User Edits
- VIAssist: Adapting Multi-modal Large Language Models for Users with Visual Impairments
- LLMSense: Harnessing LLMs for High-level Reasoning Over Spatiotemporal Sensor Traces
- HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?
- API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
- Transparent and Scrutable Recommendations Using Natural Language User Profiles
- MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
- Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security
- AppAgent: Multimodal Agents as Smartphone Users
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- VILA: On Pre-training for Visual Language Models
- OneLLM: One Framework to Align All Modalities with Language
- Explore, Select, Derive, and Recall: Augmenting LLM with Human-like Memory for Mobile Task Automation
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- EdgeFM: Leveraging Foundation Model for Open-set Learning on the Edge
- Penetrative AI: Making LLMs Comprehend the Physical World
- AutoDroid: LLM-powered Task Automation in Android
- Exploring and Characterizing Large Language Models For Embedded System Development and Debugging
- Lost in the Middle: How Language Models Use Long Contexts
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
- GPT-4 Technical Report
- Large Language Models Can Be Easily Distracted by Irrelevant Context
- A Survey on In-context Learning
- Fast Inference from Transformers via Speculative Decoding
- Learning Transferable Visual Models From Natural Language Supervision
- AdaSense: Adaptive Low-Power Sensing and Activity Recognition for Wearable Devices
- Scaling Video Analytics on Constrained Edge Nodes
- The Lumiere Project: Bayesian User Modeling for Inferring the Goals and Needs of Software Users
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Related