Small Language Models are the Future of Agentic AI
2025/06/02 by Peter Belcak, Greg Heinrich, Belcak, Peter +13 · 47 voices · 109 citations
Computer Science · #Multi-Agent Systems and Negotiation #cs.AI
paper · pdf · doi:10.48550/arxiv.2506.02153
openalex publication_date 2025/06/02 · openalex created_date 2025/10/14 · openalex updated_date 2026/07/28
Abstract
Large language models (LLMs) are often praised for exhibiting near-human performance on a wide range of tasks and valued for their ability to hold a general conversation. The rise of agentic AI systems is, however, ushering in a mass of applications in which language models perform a small number of specialized tasks repetitively and with little variation. Here we lay out the position that small language models (SLMs) are sufficiently powerful, inherently more suitable, and necessarily more economical for many invocations in agentic systems, and are therefore the future of agentic AI. Our argumentation is grounded in the current level of capabilities exhibited by SLMs, the common architectures of agentic systems, and the economy of LM deployment. We further argue that in situations where general-purpose conversational abilities are essential, heterogeneous agentic systems (i.e., agents invoking multiple different models) are the natural choice. We discuss the potential barriers for the adoption of SLMs in agentic systems and outline a general LLM-to-SLM agent conversion algorithm. Our position, formulated as a value statement, highlights the significance of the operational and economic impact even a partial shift from LLMs to SLMs is to have on the AI agent industry. We aim to stimulate the discussion on the effective use of AI resources and hope to advance the efforts to lower the costs of AI of the present day. Calling for both contributions to and critique of our position, we commit to publishing all such correspondence at https://research.nvidia.com/labs/lpr/slm-agents.
Citations
Cited by
- When Language Models Meet NeuroGraphs: Exploring Enhanced Agentic LLM Framework Towards Brain Network Analysis
- Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models
- Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis
- Enhancing Small Language Models Reasoning through Knowledge Graph Grounding
- ReMIND: Orchestrating Modular Large Language Models for Controllable Serendipity A REM-Inspired System Design for Emergent Creative Ideation
- RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation
- From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation
- Program-as-Weights: A Programming Paradigm for Fuzzy Functions
- SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification
- PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language
- LaCy: What Small Language Models Can and Should Learn is Not Just a Question of Loss
- AI Propaganda factories with language models
- It's LIT! Reliability-Optimized LLMs with Inspectable Tools
- Nested Browser-Use Learning for Agentic Information Seeking
- SANet: A Semantic-aware Agentic AI Networking Framework for Cross-layer Optimization in 6G
- Security and Privacy in Agentic AI: Grand Challenges and Future Directions
- Coordinated Networking for On-Device Agent-Augmented Real-Time Communication
- SkipOPU: An FPGA-based Overlay Processor for Large Language Models with Dynamically Allocated Computation
- An Information Theoretic Perspective on Agentic System Design
- Evaluating Small Language Models for Agentic On-Farm Decision Support Systems
- SpeakRL: Synergizing Reasoning, Speaking, and Acting in Language Models with Reinforcement Learning
- MAC: A Multi-Agent Framework for Interactive User Clarification in Multi-turn Conversations
- Generate-Then-Validate: A Novel Question Generation Approach Using Small Language Models
- SWEnergy: An Empirical Study on Energy Efficiency in Agentic Issue Resolution Frameworks with SLMs
- Reliable agent engineering should integrate machine-compatible organizational principles
- Towards Small Language Models for Security Query Generation in SOC Workflows
- David vs. Goliath: Can Small Models Win Big with Agentic AI in Hardware Design?
- Menta: A Small Language Model for On-Device Mental Health Prediction
- Small Language Models Reshape Higher Education: Courses, Textbooks, and Teaching
- STRIDE: A Systematic Framework for Selecting AI Modalities -- Agentic AI, AI Assistants, or LLM Calls
- StructuredDNA: A Bio-Physical Framework for Energy-Aware Transformer Routing
- Edge Deployment of Small Language Models, a comprehensive comparison of CPU, GPU and NPU backends
- Towards Active Synthetic Data Generation for Finetuning Language Models
- Behavior-Equivalent Token: Single-Token Replacement for Long Prompts in LLMs
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- Mitigating hallucinations and omissions in LLMs for invertible problems: An application to hardware logic design automation
- Hybrid Agentic AI and Multi-Agent Systems in Smart Manufacturing
- A Mathematical Framework for Custom Reward Functions in Job Application Evaluation using Reinforcement Learning
- A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
- AutoTool: Efficient Tool Selection for Large Language Model Agents
- Generalist Foundation Models Are Not Clinical Enough for Hospital Operations
- Enhancing LLM Code Generation Capabilities through Test-Driven Development and Code Interpreter
- Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
- Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
- Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
- Klear-AgentForge: Forging Agentic Intelligence through Posttraining Scaling
- A systematic review of relation extraction task since the emergence of Transformers
- EQ-Negotiator: Dynamic Emotional Personas Empower Small Language Models for Edge-Deployable Credit Negotiation
- A Criminology of Machines
- The Collaboration Gap
- A CPU-Centric Perspective on Agentic AI
- Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting
- Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- Validity Is What You Need
- ECVL-ROUTER: Scenario-Aware Routing for Vision-Language Models
- Training Language Models via Neural Cellular Automata
- Tongyi DeepResearch Technical Report
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- LightAgent: Mobile Agentic Foundation Models
- Accelerating Mobile Language Model via Speculative Decoding and NPU-Coordinated Execution
- Synera: Synergistic LLM Serving across Device and Cloud at Scale
- CodeEvolve: An open source evolutionary coding agent for algorithm discovery and optimization
- Toward Cybersecurity-Expert Small Language Models
- Big Reasoning with Small Models: Instruction Retrieval at Inference Time
- MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
- SVTime: Small Time Series Forecasting Models Informed by "Physics" of Large Vision Model Forecasters
- A Locally Executable AI System for Improving Preoperative Patient Communication: A Multi-Domain Clinical Evaluation
- Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
- Self-Improving LLM Agents at Test-Time
- Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the Models
- Flexible Swarm Learning May Outpace Foundation Models in Essential Tasks
- Luth: Efficient French Specialization for Small Language Models and Cross-Lingual Transfer
- Slm-mux: Orchestrating small language models for reasoning
- Dual-stage and Lightweight Patient Chart Summarization for Emergency Physicians
- GA4GC: Greener Agent for Greener Code via Multi-Objective Configuration Optimization
- RLRF: Competitive Search Agent Design via Reinforcement Learning from Ranker Feedback
- A Tale of LLMs and Induced Small Proxies: Scalable Small Language Models for Knowledge Mining
- BioVERSE: Representation Alignment of Biomedical Modalities to LLMs for Multi-Modal Reasoning
- Training Large Language Models To Reason In Parallel With Global Forking Tokens
- AgentFlux: Decoupled Fine-Tuning & Inference for On-Device Agentic Systems
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- Planner-R1: Reward Shaping Enables Efficient Agentic RL with Smaller LLMs
- On-Premise AI for the Newsroom: Evaluating Small Language Models for Investigative Document Search
- From Deferral to Learning: Online In-Context Knowledge Distillation for LLM Cascades
- MobiLLM: An Agentic AI Framework for Closed-Loop Threat Mitigation in 6G Open RANs
- EmbeddingGemma: Powerful and Lightweight Text Representations
- Federation of Agents: A Semantics-Aware Communication Fabric for Large-Scale Agentic AI
- ARE: Scaling Up Agent Environments and Evaluations
- SLM-Based Agentic AI with P-C-G: Optimized for Korean Tool Use
- An Evaluation-Centric Paradigm for Scientific Visualization Agents
- AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
- Towards General Agentic Intelligence via Environment Scaling
- Large language model-empowered next-generation computer-aided engineering
- Virtual Agent Economies
- Batch Query Processing and Optimization for Agentic Workflows
- ProST: Progressive Sub-task Training for Pareto-Optimal Multi-agent Systems Using Small Language Models
- Dynamic Sparse Attention on Mobile SoCs
- Trained Miniatures: Low cost, High Efficacy SLMs for Sales & Marketing
- Democratizing AI Development: Local LLM Deployment for India's Developer Ecosystem in the Era of Tokenized APIs
- Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities
- ArgCMV: An Argument Summarization Benchmark for the LLM-era
- Tracing Positional Bias in Financial Decision-Making: Mechanistic Insights from Qwen2.5
- Teaching LLMs to Think Mathematically: A Critical Study of Decision-Making via Optimization
- Exploring Multimodal AI Reasoning for Meteorological Forecasting from Skew-T Diagrams
- Retrieval-augmented reasoning with lean language models
- Intuition emerges in Maximum Caliber models at criticality
- EvoGraph: Hybrid Directed Graph Evolution toward Software 3.0
- Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
Discussions
- Small language models are the future of agentic AI [hn, 113 points, 45 comments]
- Small Language Models are the Future of Agentic AI [lobsters, 18 points, 9 comments]
- Billions are being thrown into AI Gigafactories for "powerful" AI/large language models. But what if it's not really that necessary? But most of the useful agentic tasks like API calls, structured out [bsky, 13 points, 1 comments]
- Small Language Models are the Future of Agentic AI arxiv.org/abs/2506.02153 これが一部で盛り上がってたけど,ちゃんと論文を読む訓練をしたことがない人が読むと不幸になるやつだと思う. [bsky, 11 points, 1 comments]
- NVIDIA posits: What if the next stage of LLMs is SLMs (smol language models)? arxiv.org/abs/2506.02153 [bsky, 8 points, 1 comments]
- arxiv.org/abs/2506.021... [bsky, 7 points, 1 comments]
- There's evidence in recent days to suggest small language models are more effective, accurate and precise . Imagine the future with language models developed by Australian companies (which are less en [bsky, 7 points, 1 comments]
- SLMs (models with fewer than 10 billion parameters, aka Small Language Models) are the future of agentic AI, according to this paper by NVIDIA. They’re sufficiently powerful for most agentic AI tasks [bsky, 6 points, 0 comments]
- Small Language Models Are the Future of Agentic AI [hn, 5 points, 0 comments]
- Small Language Models are the Future of Agentic AI arxiv.org/abs/2506.02153 #artificialintelligence #agenticAI #llms [bsky, 5 points, 0 comments]
- Ausgerechnet Forscher von NVIDIA bringen den Gedanken auf, dass es für Agentic AI wohl besser ist, auf kleine spezialisierte und effiziente Modelle zu setzen und nicht die LLM-Keule zu benutzen. [bsky, 3 points, 1 comments]
- NVIDIA now also advocating for router style switching among smaller models as more efficient for most use cases. arxiv.org/abs/2506.021... I swear to god, Silicon Valley took the most expensive, circu [bsky, 3 points, 0 comments]
- NVIDIA (empresa líder en IA generativa), en colaboración con Georgia Institute of Technology, acaban de publicar un plan de ejecución de SLMs (Small Language Models). En este proponen sustituir LLMs p [bsky, 2 points, 1 comments]
- 🤖 Small Language Models are the Future of Agentic AI > we lay out the position that small language models (SLMs) are sufficiently powerful, inherently more suitable, and necessarily more economical f [bsky, 2 points, 0 comments]
- This is aligning with some of the research I've been doing arxiv.org/pdf/2506.02153 [bsky, 2 points, 0 comments]
- As I’ve been saying for a while now: Small Language Models are the Future of Agentic AI https://arxiv.org/abs/2506.02153v1 [bsky, 1 points, 0 comments]
- arxiv.org/abs/2506.02153 [bsky, 1 points, 0 comments]
- Good to see research catching up with what math behind these models have always implied: scaling laws may favor LLMs in average generality, but data points strongly toward SSLMs for agentic use cases. [bsky, 1 points, 0 comments]
- ICYMI Small Language Models are the Future of #Agentic #AI (preprint; via #Arxiv) arxiv.org/abs/2506.02153 #SLMs (H/T the-decoder.com/heres-how-to... [bsky, 1 points, 0 comments]
- This NVIDIA position paper has a clear definition of an SLM: arxiv.org/abs/2506.02153 They consider <10B. Personally, I would not consider 13B models to be SLMs (not even 7B). They require quite a lot [bsky, 1 points, 2 comments]
- I have been experimenting with gpt-5-nano over the past week and found it to be really good at focused tasks. The trick is to partition the context into small focused chunks. Best part is it's fast an [bsky, 1 points, 0 comments]
- Paper: arxiv.org/abs/2506.02153 Project page: [bsky, 1 points, 0 comments]
- Small language models are the future of agentic AI https://arxiv.org/abs/2506.02153 https://news.ycombinator.com/item?id=44430311 [bsky, 0 points, 0 comments]
- arxiv.org/abs/2506.02153 [bsky, 0 points, 0 comments]
- Small Language Models Are the Future of Agentic AI #HackerNews https://arxiv.org/abs/2506.02153 [bsky, 0 points, 0 comments]
- https://bsky.app/profile/lobste.rs.web.brid.gy/post/3m56hxzv56mb2 [bsky, 0 points, 0 comments]
- Small Language Models are the Future of Agentic AI https://lobste.rs/s/44dgd7 #ai [bsky, 0 points, 0 comments]
- Small language models are the future of agentic AI [bsky, 0 points, 0 comments]
- 10/ As a note, even Nvidia in a 2025 research paper (arxiv.org/pdf/2506.02153) dubbed small models ‘sufficient enough’. And they keep improving. Efficiency compound growth in GPUs, hardware, etc. As s [bsky, 0 points, 2 comments]
- Genuinely surprised to discover that small models aren’t already the standard for this sort of thing. There’s a lot of interesting stuff to be done with local models specifically, and hardly anyone is [bsky, 0 points, 0 comments]
- "Our argumentation is grounded in the current level of capabilities exhibited by SLMs, the common architectures of agentic systems, and the economy of LM deployment." [bsky, 0 points, 0 comments]
- ⚡ Hackernews Top story: Small language models are the future of agentic AI [bsky, 0 points, 0 comments]
- https://bsky.app/profile/news.ycombinator.com.web.brid.gy/post/3lsv27k2knd52 [bsky, 0 points, 0 comments]
- SLMはより扱いやすく、デプロイ可能性が高い セキュリティ、計算コスト、応答速度、制御性といった現実的な制約の下では、LLMよりSLMの方が扱いやすい。LLMは「ブラックボックス化」が進んでおり、予測不可能な振る舞いをすることがあるが、SLMはチューニングしやすく透明性が高い。 Small Language Models are the Future of Agentic AI arxiv.org [bsky, 0 points, 0 comments]
- arxiv.org/pdf/2506.02153 [bsky, 0 points, 0 comments]
- Many think so! arxiv.org/abs/2506.02153 [bsky, 0 points, 0 comments]
- This has been my bet for a while arxiv.org/pdf/2506.02153 [bsky, 0 points, 0 comments]
- [2506.02153] Small Language Models are the Future of Agentic AI arxiv.org/abs/2506.02153 [bsky, 0 points, 0 comments]
- Малые языковые модели - это будущее агентного ИИ #ai #news [bsky, 0 points, 0 comments]
- Small language models are the future of agentic AI #ai #news [bsky, 0 points, 0 comments]
- The next AI tool you use might not need a supercomputer—or even the cloud. It could run fast and locally on your phone or PC. Small language models are catching up, and they’re cheaper, leaner and app [bsky, 0 points, 0 comments]
- 5/ That hybrid setup keeps costs near zero per call and still gives you power when you need it. This is something I’ve started testing in my projects, and it seems like a very interesting direction. P [bsky, 0 points, 0 comments]
- Small language models are the future of agentic AI https://arxiv.org/abs/2506.02153 (https://news.ycombinator.com/item?id=44430311) [bsky, 0 points, 0 comments]
- Small language models are the future of agentic AI https://arxiv.org/abs/2506.02153 (https://news.ycombinator.com/item?id=44430311) [bsky, 0 points, 0 comments]
- Small Language Models Are the Future of Agentic AI https://arxiv.org/abs/2506.02153 [bsky, 0 points, 0 comments]
- → Most agentic sub-tasks are narrow and repetitive → LLMs are catastrophically overqualified for 80% of those tasks → SLMs run faster, cheaper, and with far less energy We're using a sledgehammer to p [bsky, 0 points, 0 comments]
- The NVIDIA SLM paper from Belcak and Heinrich puts the cost gap at roughly 10–30× cheaper per call than a 70B+ model. When iterations are that cheap, you can afford to be wasteful in the right directi [bsky, 0 points, 1 comments]
Related