HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
2023/03/30 by Yongliang Shen, Shen, Yongliang, Kaitao Song +10 · 2 voices · 313 citations
Computer Science · Engineering · #Artificial intelligence #Computer science #Face (sociological concept) #Ferroelectric and Negative Capacitance Devices #Human–computer interaction #Interface (matter) #Linguistics #Modalities #Multimodal Machine Learning Applications #Natural language processing #Task (project management) #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2303.17580
openalex publication_date 2023/03/30 · openalex created_date 2023/04/05 · openalex updated_date 2026/07/28
Abstract
Solving complicated AI tasks with different domains and modalities is a key step toward artificial general intelligence. While there are numerous AI models available for various domains and modalities, they cannot handle complicated AI tasks autonomously. Considering large language models (LLMs) have exhibited exceptional abilities in language understanding, generation, interaction, and reasoning, we advocate that LLMs could act as a controller to manage existing AI models to solve complicated AI tasks, with language serving as a generic interface to empower this. Based on this philosophy, we present HuggingGPT, an LLM-powered agent that leverages LLMs (e.g., ChatGPT) to connect various AI models in machine learning communities (e.g., Hugging Face) to solve AI tasks. Specifically, we use ChatGPT to conduct task planning when receiving a user request, select models according to their function descriptions available in Hugging Face, execute each subtask with the selected AI model, and summarize the response according to the execution results. By leveraging the strong language capability of ChatGPT and abundant AI models in Hugging Face, HuggingGPT can tackle a wide range of sophisticated AI tasks spanning different modalities and domains and achieve impressive results in language, vision, speech, and other challenging tasks, which paves a new way towards the realization of artificial general intelligence.
Cited by
- SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision
- Encoding Invisible Causation for Bridge Diagnostic Agents: Triple-Guided Retrieval-Augmented Fine-Tuning with QLoRA
- Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
- Personalized Recommendation Tool Learning via Autonomous Language Agents
- Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development
- Operational Hallucination and Safety Drift in AI Agents
- RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control
- SkillRouter: Skill Routing for LLM Agents at Scale
- Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- Knowledge-Centric Agents for Workflow Generation in ComfyUI
- PhotoAgent: Exploratory Visual Aesthetic Planning with Large Vision Models
- AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
- AgentWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
- StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
- Position: Modular Memory is the Key to Continual Learning Agents
- Smart Microscopy: Current Implementations and a Roadmap for Interoperability
- Coral Protocol: Open Infrastructure Connecting The Internet of Agents
- It's LIT! Reliability-Optimized LLMs with Inspectable Tools
- Nested Browser-Use Learning for Agentic Information Seeking
- Video Understanding: From Geometry and Semantics to Unified Models
- MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
- CountGD++: Generalized Prompting for Open-World Counting
- LLMTM: Benchmarking and Optimizing LLMs for Temporal Motif Analysis in Dynamic Graphs
- A Survey on Large Language Model-based Agents for Statistics and Data Science
- SymStep: Symbolic Step Verification for Logical Reasoning
- VeraGrid-Agent: Tool-Augmented LLMs for Distribution Optimal Power Flow at the Grid Edge
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
- Reason Before You Retrieve: Agentic Planning for Multi-modal RAG
- MoRAgent: Parameter Efficient Agent Tuning with Mixture-of-Roles
- A Plan Reuse Mechanism for LLM-Driven Agent
- RoboSafe: Safeguarding Embodied Agents via Executable Safety Logic
- ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- LLMCache: Layer-Wise Caching Strategies for Accelerated Reuse in Transformer Inference
- SynthSeg-Agents: Multi-Agent Synthetic Data Generation for Zero-Shot Weakly Supervised Semantic Segmentation
- Evidence-Driven Decision Support for AI Model Selection in Research Software Engineering
- AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning
- AgentSHAP: Interpreting LLM Agent Tool Importance with Monte Carlo Shapley Value Estimation
- AgentBalance: Backbone-then-Topology Design for Cost-Effective Multi-Agent Systems under Budget Constraints
- Asynchronous Reasoning: Training-Free Interactive Thinking LLMs
- AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
- Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- Poodle: Seamlessly Scaling Down Large Language Models with Just-in-Time Model Replacement
- RevoNAD: Reflective Evolutionary Exploration for Neural Architecture Design
- Embodied Co-Design for Rapidly Evolving Agents: Taxonomy, Frontiers, and Challenges
- PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
- On Group Relative Policy Optimization Collapse in Agent Search: The Lazy Likelihood-Displacement
- Measuring Agents in Production
- SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts
- STRIDE: A Systematic Framework for Selecting AI Modalities -- Agentic AI, AI Assistants, or LLM Calls
- COACH: Collaborative Agents for Contextual Highlighting -- A Multi-Agent Framework for Sports Video Analysis
- Energy-Aware Data-Driven Model Selection in LLM-Orchestrated AI Systems
- GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
- EWE: An Agentic Framework for Extreme Weather Analysis
- Subgoal Graph-Augmented Planning for LLM-Guided Open-World Reinforcement Learning
- The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
- Complex QA and language models hybrid architectures, Survey
- VICoT-Agent: A Vision-Interleaved Chain-of-Thought Framework for Interpretable Multimodal Reasoning and Scalable Remote Sensing Analysis
- Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query Optimization
- HuggingR4: A Progressive Reasoning Framework for Discovering Optimal Model Companions
- ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
- M3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
- ARISE: Agentic Rubric-Guided Iterative Survey Engine for Automated Scholarly Paper Generation
- AutoBackdoor: Automating Backdoor Attacks via LLM Agents
- What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
- AutoTool: Efficient Tool Selection for Large Language Model Agents
- Ghost in the Minecraft: Generally Capable Agents for Open-World Environments via Large Language Models with Text-based Knowledge and Memory
- Draft and Refine with Visual Experts
- Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
- LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
- OSGym: Scalable OS Infra for Computer Use Agents
- UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
- Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
- A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archives
- PublicAgent: Multi-Agent Design Principles From an LLM-Based Open Data Analysis Framework
- ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planning
- OceanAI: A Conversational Platform for Accurate, Transparent, Near-Real-Time Oceanographic Insights
- GraphChain: Large Language Models for Large-scale Graph Analysis via Tool Chaining
- FlowMesh: A Service Fabric for Composable LLM Workflows
- Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
- Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
- LLM-TSFD: An industrial time series human-in-the-loop fault diagnosis method based on a large language model
- AI literacy as a core component of AI education
- Flows: Building Blocks of Reasoning and Collaborating AI
- ScaleCall -- Agentic Tool Calling at Scale for Fintech: Challenges, Methods, and Deployment Insights
- Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
- Integrating Machine Learning into Belief-Desire-Intention Agents: Current Advances and Open Challenges
- AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
- Adaptive Minds: Empowering Agents with LoRA-as-Tools
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
- ToolPRM: Fine-Grained Inference Scaling of Structured Outputs for Function Calling
- LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?
- Disaster Management in the Era of Agentic AI Systems: A Vision for Collective Human-Machine Intelligence for Augmented Resilience
- NetMCP: Network-Aware Model Context Protocol Platform for LLM Capability Extension
- A Survey on Evaluation of Large Language Models
- GOAT: A Training Framework for Goal-Oriented Agent with Tools
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents
- What Slows Down FMware Development? An Empirical Study of Developer Challenges and Resolution Times
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Fundamentals of Building Autonomous LLM Agents
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- MoA-VR: A Mixture-of-Agents System Towards All-in-One Video Restoration
- A2Search: Ambiguity-Aware Question Answering with Reinforcement Learning
- Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions
- FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipeline
- ToolMem: Enhancing Multimodal Agents with Learnable Tool Capability Memory
- Adaptive Tool Generation with Models as Tools and Reinforcement Learning
- Natural Language Edge Labelling: Decoupling Intent from Execution in Structured LM Reasoning
- AlphaApollo: A System for Deep Agentic Reasoning
- Self-Planning Code Generation with Large Language Models
- AutoMaAS: Self-Evolving Multi-Agent Architecture Search for Large Language Models
- AgenticRAG: Tool-Augmented Foundation Models for Zero-Shot Explainable Recommender Systems
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- XR Blocks: Accelerating Human-centered AI + XR Innovation
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Agentic Services Computing
- Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
- AutoEP: LLMs-Driven Automation of Hyperparameter Evolution for Metaheuristic Algorithms
- Understanding and Enhancing the Planning Capability of Language Models via Multi-Token Prediction
- Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
- CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
- VC-Agent: An Interactive Agent for Customized Video Dataset Collection
- CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineering
- PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
- Federation of Agents: A Semantics-Aware Communication Fabric for Large-Scale Agentic AI
- Online-Optimized RAG for Tool Use and Function Calling
- SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
- Auditing Emergent LLM-Agent Collaboration through Cooperation-Obligation Coupling
- Can LLMs Reason Over Non-Text Modalities in a Training-Free Manner? A Case Study with In-Context Representation Learning
- SignalLLM: A General-Purpose LLM Agent Framework for Automated Signal Processing
- Evaluating the Limitations of Local LLMs in Solving Complex Programming Challenges
- Foam-Agent 2.0: An End-to-End Composable Multi-Agent Framework for Automating CFD Simulation in OpenFOAM
- PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
- UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
- Jupiter: Enhancing LLM Data Analysis Capabilities via Notebook and Inference-Time Value-Guided Search
- Code2MCP: Transforming Code Repositories into MCP Services
- Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- Guideline-Consistent Segmentation via Multi-Agent Refinement
- VISP: Volatility Informed Stochastic Projection for Adaptive Regularization
- Batch Query Processing and Optimization for Agentic Workflows
- Dynamic Speculative Agent Planning
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- One VLM, Two Roles: Stage-Wise Routing and Specialty-Level Deployment for Clinical Workflows
- Provable Benefits of In-Tool Learning for Large Language Models
- Network-Level Prompt and Trait Leakage in Local Research Agents
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- AppAgent-Pro: A Proactive GUI Agent System for Multidomain Information Integration and User Assistance
- Open-Universe Assistance Games
- GTool: Graph Enhanced Tool Planning with Large Language Model
- TASER: Table Agents for Schema-guided Extraction and Recommendation
- SSRL: Self-Search Reinforcement Learning
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education
- KonfAI: A Modular and Fully Configurable Framework for Deep Learning in Medical Imaging
- DevNous: An LLM-Based Multi-Agent System for Grounding IT Project Management in Unstructured Conversation
- Large Language Models in the Data Science Lifecycle: A Systematic Mapping Study
- Training-Free Multimodal Large Language Model Orchestration
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- Parallelism Meets Adaptiveness: Scalable Documents Understanding in Multi-Agent LLM Systems
- Unified Tool Integration for LLMs: A Protocol-Agnostic Approach to Function Calling
- AQUAH: Automatic Quantification and Unified Agent in Hydrology
- CABENCH: Benchmarking Composable AI for Solving Complex Tasks through Composing Ready-to-Use Models
- HealthFlow: A Self-Evolving AI Agent with Meta Planning for Autonomous Healthcare Research
- Interleaved LLM and Motion Planning for Generalized Multi-Object Collection in Large Scene Graphs
- Agentic AI for autonomous anomaly management in complex systems
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
- MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
- Graph-Augmented Large Language Model Agents: Current Progress and Future Prospects
- T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation
- Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
- RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
- IM-Chat: A Multi-agent LLM Framework Integrating Tool-Calling and Diffusion Modeling for Knowledge Transfer in Injection Molding Industry
- Advancing Responsible Innovation in Agentic AI: A study of Ethical Frameworks for Household Automation
- MindFlow+: A Self-Evolving Agent for E-Commerce Customer Service
- GUI-G2: Gaussian Reward Modeling for GUI Grounding
- FastLongSpeech: Enhancing Large Speech-Language Models for Efficient Long-Speech Processing
- A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents
- SoilNet: A multimodal multitask model for hierarchical classification of soil horizons
- Adaptive Multi-Agent Reasoning via Automated Workflow Generation
- FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
- Aime: Towards Fully-Autonomous Multi-Agent Framework
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
- Teaching Physical Awareness to LLMs through Sounds
- Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
- SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition
- ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs
- PyVision: Agentic Vision with Dynamic Tooling
- TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
- MARBLE: A Multi-Agent Rule-Based LLM Reasoning Engine for Accident Severity Prediction
- Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks
- Scaling LLM Planning: NL2FLOW for Parametric Problem Generation and Rigorous Evaluation
- The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems
- Generative Exaggeration in LLM Social Agents: Consistency, Bias, and Toxicity
- MassTool: A Multi-Task Search-Based Tool Retrieval Framework for Large Language Models
- iPanda: An LLM-based Agent for Automated Conformance Testing of Communication Protocols
- AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks
- SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
- A Modular Multitask Reasoning Framework Integrating Spatio-temporal Models and LLMs
- CoMind: Towards Community-Driven Agents for Machine Learning Engineering
- A Survey of Behavior Foundation Model: Next-Generation Whole-Body Control System of Humanoid Robots
- SAFEFLOW: A Principled Protocol for Trustworthy and Transactional Autonomous Agent Systems
- KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models
- NaviAgent: Graph-Driven Bilevel Planning for Scalable Tool Orchestration
- Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models
- GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System
- ReFrame: Rectification Framework for Image Explaining Architectures
- General-Purpose Robotic Navigation via LVLM-Orchestrated Perception, Reasoning, and Acting
- PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases
- Arch-Router: Aligning LLM Routing with Human Preferences
- Intelligent Assistants for the Semiconductor Failure Analysis with LLM-Based Planning Agents
- Efficient Serving of LLM Applications with Probabilistic Demand Modeling
- StorySage: Conversational Autobiography Writing Powered by a Multi-Agent Framework
- Translating Federated Learning Algorithms in Python into CSP Processes Using ChatGPT
- LocationReasoner: Evaluating LLMs on Real-World Site Selection Reasoning
- Tiered Agentic Oversight: A Hierarchical Multi-Agent System for Healthcare Safety
- Eliciting Reasoning in Language Models with Cognitive Tools
- PE-MA: Parameter-Efficient Co-Evolution of Multi-Agent Systems
- Interaction, Process, Infrastructure: A Unified Framework for Human-Agent Collaboration
- Can Theoretical Physics Research Benefit from Language Agents?
- Towards Next-Generation Intelligent Maintenance: Collaborative Fusion of Large and Small Models
- AgentSwift: Efficient LLM Agent Design via Value-guided Hierarchical Search
- Gen-n-Val: Agentic Image Data Generation and Validation
- Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling
- Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
- TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems
- Optimization Problem Solving Can Transition to Evolutionary Agentic Workflows
- Understanding Physical Properties of Unseen Deformable Objects by Leveraging Large Language Models and Robot Actions
- How Far Are We from Generating Missing Modalities with Foundation Models?
- AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance
- HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
- The Future of Continual Learning in the Era of Foundation Models: Three Key Directions
- MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
- STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset
- MCP-Zero: Active Tool Discovery for Autonomous LLM Agents
- Simple Prompt Injection Attacks Can Leak Personal Data Observed by LLM Agents During Task Execution
- MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility
- SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
- From Glue-Code to Protocols: A Critical Analysis of A2A and MCP Integration for Scalable Agent Systems
- PhotoArtAgent: Intelligent Photo Retouching with Language Model-Based Artist Agents
- ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks
- Agent-UniRAG: A Trainable Open-Source LLM Agent Framework for Unified Retrieval-Augmented Generation Systems
- Universal Visuo-Tactile Video Understanding for Embodied Interaction
- MLE-STAR: Machine Learning Engineering Agent via Search and Targeted Refinement
- Agentic 3D Scene Generation with Spatially Contextualized VLMs
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play
- MetaMind: Modeling Human Social Thoughts with Metacognitive Multi-Agent Systems
- CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
- LiteCUA: Computer as MCP Server for Computer-Use Agent on AIOS
- AI-Researcher: Autonomous Scientific Innovation
- BEDI: A Comprehensive Benchmark for Evaluating Embodied Agents on UAVs
- The Real Barrier to LLM Agent Usability is Agentic ROI
- Benchmarking Recommendation, Classification, and Tracing Based on Hugging Face Knowledge Graph
- T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
- Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
- VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
- Position: Agentic Systems Constitute a Key Component of Next-Generation Intelligent Image Processing
- Adaptive Plan-Execute Framework for Smart Contract Security Auditing
- Structured Agent Distillation for Large Language Model
- InfiGFusion: Graph-on-Logits Distillation via Efficient Gromov-Wasserstein for Model Fusion
- Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
- Understanding Complexity in VideoQA via Visual Program Generation
- MADE: Belief-Driven Dual-Agent Coordination for Autonomous Model Deployment
- MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging
- DiffuseAgent-MI: Distributionally-Grounded,Tool-Integrated Self-Evolving Agents for Faithful Visual Reasoning
- Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction
- CartoAgent: a multimodal large language model-powered multi-agent cartographic framework for map style transfer and evaluation
- 14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon
- Strategy-Augmented Planning for Large Language Models via Opponent Exploitation
- Agent-as-a-Service based on Agent Network
- TUMS: Enhancing Tool-use Abilities of LLMs with Multi-structure Handlers
- DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
- FitText: Evolving Agent Tool Ecologies via Memetic Retrieval
- Putting It All into Context: Simplifying Agents with LCLMs
- Internet of Agents: Fundamentals, Applications, and Challenges
- RideAgent: An LLM-Enhanced Optimization Framework for Automated Taxi Fleet Operations
- From OSS to Open Source AI: an Exploratory Study of Collaborative Development Paradigm Divergence
- AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference
- WirelessAgent: Large Language Model Agents for Intelligent Wireless Networks
- ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents
- Holos: A Web-Scale LLM-Based Multi-Agent System for the Agentic Web
- UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
- Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
- Separating Intelligence from Execution: A Workflow Engine for the Model Context Protocol
- Benchmark Test-Time Scaling of General LLM Agents
- FactorMiner: A Self-Evolving Agent with Skills and Experience Memory for Financial Alpha Discovery
- Agent libOS: A Runtime Substrate for Capability-Controlled Self-Evolving LLM Agents
- Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI
- NGENT: Next-Generation AI Agents Must Integrate Multi-Domain Abilities to Achieve Artificial General Intelligence
- Self-Distilled Agentic Reinforcement Learning
- An Empirical Study on the Capability of LLMs in Decomposing Bug Reports
- PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement
- SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
- Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows
- Harness as an Asset: Enforcing Determinism via the Convergent AI Agent Framework (CAAF)
- AgentGate: A Lightweight Structured Routing Engine for the Internet of Agents
- When Single-Agent with Skills Replace Multi-Agent Systems and When They Fail
- ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- Symbolic Representation for Any-to-Any Generative Tasks
- Blockchain Empowered Trustworthy Agent Networks: Foundations, Taxonomy, and Future Directions
- TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
- A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms
- Empowering AI to Generate Better AI Code: Guided Generation of Deep Learning Projects with LLMs
- DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
Discussions
Related