DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
2023/10/05 by Omar Khattab, Arnav Singhvi, Khattab, Omar +23 · 122 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Software Engineering Research
paper · pdf · doi:10.48550/arxiv.2310.03714
Abstract
The ML community is rapidly exploring techniques for prompting language models (LMs) and for stacking them into pipelines that solve complex tasks. Unfortunately, existing LM pipelines are typically implemented using hard-coded "prompt templates", i.e. lengthy strings discovered via trial and error. Toward a more systematic approach for developing and optimizing LM pipelines, we introduce DSPy, a programming model that abstracts LM pipelines as text transformation graphs, i.e. imperative computational graphs where LMs are invoked through declarative modules. DSPy modules are parameterized, meaning they can learn (by creating and collecting demonstrations) how to apply compositions of prompting, finetuning, augmentation, and reasoning techniques. We design a compiler that will optimize any DSPy pipeline to maximize a given metric. We conduct two case studies, showing that succinct DSPy programs can express and optimize sophisticated LM pipelines that reason about math word problems, tackle multi-hop retrieval, answer complex questions, and control agent loops. Within minutes of compiling, a few lines of DSPy allow GPT-3.5 and llama2-13b-chat to self-bootstrap pipelines that outperform standard few-shot prompting (generally by over 25% and 65%, respectively) and pipelines with expert-created demonstrations (by up to 5-46% and 16-40%, respectively). On top of that, DSPy programs compiled to open and relatively small LMs like 770M-parameter T5 and llama2-13b-chat are competitive with approaches that rely on expert-written prompt chains for proprietary GPT-3.5. DSPy is available at https://github.com/stanfordnlp/dspy
Cited by
- Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
- HiFi-RAG: Hierarchical Content Filtering and Two-Pass Generation for Open-Domain RAG
- Moral Hazard in Multi-Agent Language Models
- A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain
- Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents
- Imprompt: A Language Framework for Prompt Programming
- Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning
- FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
- ContextLeak: Auditing Leakage in Private In-Context Learning Methods
- SCOPE: Prompt Evolution for Enhancing Agent Effectiveness
- The Meta-Prompting Protocol: Orchestrating LLMs via Adversarial Feedback Loops
- Sharing State Between Prompts and Programs
- Verification-Guided Context Optimization for Tool Calling via Hierarchical LLMs-as-Editors
- Textual Gradients are a Flawed Metaphor for Automatic Prompt Optimization
- FutureWeaver: Planning Test-Time Compute for Multi-Agent Systems with Modularized Collaboration
- Exploring LLMs for Scientific Information Extraction Using The SciEx Framework
- Small Language Models Can Use Nuanced Reasoning For Health Science Research Classification: A Microbial-Oncogenesis Case Study
- Lumos: Let there be Language Model System Certification
- In-Context Distillation with Self-Consistency Cascades: A Simple, Training-Free Way to Reduce LLM Agent Costs
- Automated Risk-of-Bias Assessment of Randomized Controlled Trials: A First Look at a GEPA-trained Programmatic Prompting Framework
- ThetaEvolve: Test-time Learning on Open Problems
- Aragog: Just-in-Time Model Routing for Scalable Serving of Agentic Workflows
- Structured Prompting Enables More Robust Evaluation of Language Models
- AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
- Prompt Less, Smile More: MTP with Semantic Engineering in Lieu of Prompt Engineering
- Prompt Optimization as a State-Space Search Problem
- LangMark: A Multilingual Dataset for Automatic Post-Editing
- WER is Unaware: Assessing How ASR Errors Distort Clinical Understanding in Patient Facing Dialogue
- PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization
- ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process Routing
- Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues
- RAGSmith: A Framework for Finding the Optimal Composition of Retrieval-Augmented Generation Methods Across Datasets
- Preference Learning from Physics-Based Feedback: Tuning Language Models to Design BCC/B2 Superalloys
- Prompt Triage: Structured Optimization Enhances Vision-Language Model Performance on Medical Imaging Benchmarks
- PerspAct: Enhancing LLM Situated Collaboration Skills through Perspective Taking and Active Vision
- Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
- A Picture is Worth a Thousand (Correct) Captions: A Vision-Guided Judge-Corrector System for Multimodal Machine Translation
- RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
- Batch Prompting Suppresses Overthinking Reasoning Under Constraint: How Batch Prompting Suppresses Overthinking in Reasoning Models
- Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything
- Continual Learning, Not Training: Online Adaptation For Agents
- Can Language Models Go Beyond Coding? Assessing the Capability of Language Models to Build Real-World Systems
- On Selecting Few-Shot Examples for LLM-based Code Vulnerability Detection
- Decomposition-Enhanced Training for Post-Hoc Attributions In Language Models
- TextualVerifier: Verify TextGrad Step-by-Step
- Rethinking the Evaluation of Harness Evolution for Agents
- Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
- IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations
- SAGE: Stochastic Prompt Optimization via Agent-Guided Exploration
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- Compiler.next: A Search-Based Compiler to Power the AI-Native Future of Software Engineering
- Thought Communication in Multiagent Collaboration
- Code-enabled language models can outperform reasoning models on diverse tasks
- Prompt Decorators: A Declarative and Composable Syntax for Reasoning, Formatting, and Control in LLMs
- VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
- PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits
- SAVANT: Semantic Analysis with Vision-Augmented Anomaly deTection
- Lexo: Eliminating Stealthy Supply-Chain Attacks via LLM-Assisted Program Regeneration
- Natural Language Tools: A Natural Language Approach to Tool Calling In Large Language Agents
- Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
- SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
- From Craft to Constitution: A Governance-First Paradigm for Principled Agent Engineering
- Failure-Driven Workflow Refinement
- RoboPhD: Evolving Diverse Complex Agents Under Tight Evaluation Budgets
- Meta-Harness: End-to-End Optimization of Model Harnesses
- How can we assess human-agent interactions? Case studies in software agent design
- Agentic Systems in Radiology: Design, Applications, Evaluation, and Challenges
- Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
- Efficient Tree-Structured Deep Research with Adaptive Resource Allocation
- ARM: Discovering Agentic Reasoning Modules for Generalizable Multi-Agent Systems
- Slm-mux: Orchestrating small language models for reasoning
- JSON Whisperer: Efficient JSON Editing with LLMs
- TimeSeriesScientist: A General-Purpose AI Agent for Time Series Analysis
- HLTCOE at TREC 2024 NeuCLIR Track
- ACT: Agentic Classification Tree
- The Hunger Game Debate: On the Emergence of Over-Competition in Multi-Agent Systems
- DyFlow: Dynamic Workflow Framework for Agentic Reasoning
- Towards Reliable Generation of Executable Workflows by Foundation Models
- Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE
- A Large-Scale Dataset and Citation Intent Classification in Turkish with LLMs
- AgentPack: A Dataset of Code Changes, Co-Authored by Agents and Humans
- A State-of-the-Art SQL Reasoning Model using RLVR
- Eigen-1: Adaptive Multi-Agent Refinement with Monitor-Based RAG for Scientific Reasoning
- Structuring Collective Action with LLM-Guided Evolution: From Ill-Structured Problems to Executable Heuristics
- DELM: a Python toolkit for Data Extraction with Language Models
- Human-AI Narrative Synthesis to Foster Shared Understanding in Civic Decision-Making
- Score the Steps, Not Just the Goal: VLM-Based Subgoal Evaluation for Robotic Manipulation
- What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
- Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
- Cybersecurity Detection Classification with Reasoning-enabled Language Models
- AutoMem: Automated Learning of Memory as a Cognitive Skill
- Bringing Pedagogy into Focus: Evaluating Virtual Teaching Assistants' Question-Answering in Asynchronous Learning Environments
- Integrating Text and Time-Series into (Large) Language Models to Predict Medical Outcomes
- Do Code Semantics Help? A Comprehensive Study on Execution Trace-Based Information for Code Large Language Models
- Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent Workflows
- A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment
- PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability
- A Composable Agentic System for Automated Visual Data Reporting
- Maestro: Joint Graph & Config Optimization for Reliable AI Agents
- AgenTracer: Who Is Inducing Failure in the LLM Agentic Systems?
- Batch Query Processing and Optimization for Agentic Workflows
- Towards Temporal Knowledge-Base Creation for Fine-Grained Opinion Analysis with Language Models
- SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization
- CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering
- Type-Compliant Adaptation Cascades: Adapting Programmatic LM Workflows to Data
- Why Synthetic Isn't Real Yet: A Diagnostic Framework for Contact Center Dialogue Generation
- InPars+: Supercharging Synthetic Data Generation for Information Retrieval Systems
- Type-Driven Prompt Programming: From Typed Interfaces to a Calculus of Constraints
- Searching for Privacy Risks in LLM Agents via Simulation
- PASS: Probabilistic Agentic Supernet Sampling for Interpretable and Adaptive Chest X-Ray Reasoning
- LingVarBench: Benchmarking LLM for Automated Named Entity Recognition in Structured Synthetic Spoken Transcriptions
- LLM Empowered Prototype Learning for Zero and Few-Shot Tasks on Tabular Data
- GreenTEA: Gradient Descent with Topic-modeling and Evolutionary Auto-prompting
- When the Domain Expert Has No Time and the LLM Developer Has No Clinical Expertise: Real-World Lessons from LLM Co-Design in a Safety-Net Hospital
- When Prompt Engineering Meets Software Engineering: CNL-P as Natural and Robust "APIs'' for Human-AI Interaction
- Making Prompts First-Class Citizens for Adaptive LLM Pipelines
- A Survey on Agent Workflow -- Status and Future
- Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates
- CompoST: A Benchmark for Analyzing the Ability of LLMs To Compositionally Interpret Questions in a QALD Setting
- Auto: The AGI Compiler
- Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
- Expert-aided causal discovery of ancestral graphs
Related