Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Prompt Optimization with DSPy

By Anup Rai5 min readReviewed September 2026

DSPy is a framework for expressing language-model programs and optimizing them against examples and a metric. It separates the program's input/output contracts and control flow from some of the instructions and demonstrations used to implement them. An optimizer searches for a better configuration; it does not prove that it has found the best possible prompt.

The original DSPy paper describes compiling declarative language-model calls into improved pipelines. The framework is one option for systematic optimization, not a requirement for every LLM application. Khattab et al..

Learn the four building blocks

Element Meaning Example
Signature Declares task inputs and outputs Ticket text → routing label
Module Implements one or more model calls and program behavior A classifier followed by evidence validation
Metric Scores whether the program achieved the task Correct label and a valid source excerpt
Optimizer Searches instructions, examples or other supported parameters Select demonstrations using development feedback

A signature describes a contract. It does not supply retrieval, authorization or a correct algorithm merely because its description says “multi-hop reasoning.” If the task requires several searches, the program must actually perform them.

Start with a program that can be evaluated

For the support-routing task, define:

  1. Inputs: ticket text and the versioned routing rubric.
  2. Outputs: one allowed label and an exact supporting excerpt, when present.
  3. Control flow: predict, validate the output contract, then return a routing proposal.
  4. Evaluation: correct label, excerpt grounded in the ticket, and explicit handling of ambiguous cases.
  5. Boundaries: no account modifications or private cross-tenant examples.

The core task contract can be expressed without any framework:

route(ticket_text, rubric_version) -> label, evidence_excerpt

Allowed labels: cancel, refund, multiple, other
Evidence excerpt: exact substring of the input, or empty when appropriate
Side effects: none

Implement this as a minimal baseline before adding an optimizer. A framework cannot make an undefined label policy measurable.

Choose an optimizer for the actual problem

Optimizer family What it changes What to inspect
Labeled or bootstrapped few-shot Demonstrations; bootstrapping can use successful program traces Example correctness and representativeness
MIPROv2 Instructions and demonstrations, using candidate evaluation and search Search budget, development-set overfitting
GEPA Prompt candidates informed by reflection on trajectories and feedback Quality and privacy of feedback; objective gaming
BootstrapFinetune Model weights using generated/selected training traces Training quality, supported model and deployment costs

These distinctions reflect the current DSPy optimizer documentation. Prompt optimization usually leaves the base model weights unchanged. Some DSPy optimizers explicitly perform fine-tuning, so “DSPy never changes weights” would also be wrong.

“Compile” in this context means producing an optimized program configuration. It does not imply a formal correctness proof or necessarily gradient descent. Avoid assuming that a prompt is literally a differentiable neural-network weight.

Keep evaluation independent of selection

Architecture / visual model
flowchart LR T[Training examples and traces] --> O[Generate candidate programs] D[Development cases and metric] --> O O --> S[Select and freeze candidate] S --> H[Evaluate on untouched test set] H --> R[Release or investigate failures]
Read diagram source
flowchart LR
    T[Training examples and traces] --> O[Generate candidate programs]
    D[Development cases and metric] --> O
    O --> S[Select and freeze candidate]
    S --> H[Evaluate on untouched test set]
    H --> R[Release or investigate failures]

Optimizer APIs differ in how they use training and validation inputs. Whatever the interface, keep a final test set outside the search. Repeatedly choosing prompts based on that test set turns it into another development set.

Split related examples by the relevant unit: customer, document family, time period or task instance. Randomly splitting near-duplicate tickets can exaggerate generalization. Review rare labels, adversarial instructions and inputs requiring abstention separately.

Design a metric the system cannot cheaply exploit

Exact match is useful for a fixed routing label. It is less appropriate for an open-ended answer with several correct phrasings. A model judge can assess a rubric, but needs calibration against human-reviewed cases and checks for position, style and length bias.

Use hard acceptance conditions for properties that cannot be traded away. An aggregate score should not allow a modest accuracy improvement to compensate for unauthorized data access. Evaluate those boundaries independently of the prompt optimizer.

For the classifier, report label accuracy, per-label errors, evidence validity, invalid-output rate and latency. If the objective rewards only accuracy on common cases, the optimizer may find a short prompt that performs badly on rare categories.

Budget the search and the deployed program

An illustrative search with 100 candidate evaluations, 50 cases per evaluation and two model calls per case uses 100 × 50 × 2 = 10,000 task-model calls. Teacher generation, reflection, retries and final evaluation add work. Some algorithms reuse results or evaluate subsets; inspect actual accounting rather than assuming this full schedule.

There are two costs to justify: optimization cost and steady-state serving cost. An optimized program with longer demonstrations may improve quality while increasing every future request's input tokens. Compare total lifecycle cost and the cost per successful task.

Release and model changes

Version the program, optimized instructions/examples, model, adapters, generation settings, dataset and metric. Preserve the baseline and rollback path. Monitor real traffic for distribution changes and investigate new failures before searching again.

For a model upgrade, evaluate the existing program first. Re-optimize if the results justify it. A model change does not logically require rewriting every prompt, and recompilation does not automatically restore previous quality.

Interview practice

Q1: What problem does DSPy solve?

It provides structure for composing model programs and searching configurations against a metric. It reduces ad hoc prompt selection when there are representative examples and a meaningful evaluation. It does not remove the need to define the task, data access or acceptance criteria.

Q2: Does a signature implement a multi-hop retrieval system?

No. It describes what a call receives and returns. The program must implement retrieval, intermediate state and stopping behavior. A descriptive class name cannot create missing control flow.

Q3: How does MIPROv2 differ from GEPA at a high level?

MIPROv2 searches instruction and demonstration candidates using performance feedback. GEPA uses reflective feedback on execution trajectories to propose prompt changes. I would select based on the available metric/feedback, search budget and measured results rather than declaring one universally superior.

Q4: Is prompt optimization fine-tuning?

Not when it changes instructions and examples only. Fine-tuning changes model parameters. DSPy supports both kinds of optimization through different components, so I would specify exactly what the chosen optimizer changes.

Q5: Why might the highest-scoring candidate be a poor release?

It may overfit the development cases, exploit a weak judge, increase serving cost or fail an important category. Freeze the candidate, use an untouched test set, inspect failure slices and apply separate operational and security acceptance criteria.

Q6: What if there are only thirty reviewed examples?

Start with a simple baseline and a modest search. Preserve independent checks, inspect every failure and collect more representative data. A large automated search against a tiny dataset can select noise. More search is not a substitute for better evidence.

Final notes

Recall card: Contract → program → metric → search → independent test. Automated prompt optimization is an empirical development process, with versioning and release controls like the rest of the application.

Related: few-shot learning, structured generation, fine-tuning.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Structured Generation
NEXT LESSONPrompt Injection and Defense →

Explore the diagram