Enterprise DNA
Guide Intermediate General

DSPy Tutorial for Better LLM Applications

Build a DSPy ticket-routing application and learn signatures, modules, optimizers, evaluation, and practical prompt improvement.

Sam McKay |
DSPy Tutorial for Better LLM Applications

DSPy is a Python framework for building language-model applications that can improve from examples instead of relying on manually maintained prompt text. In this tutorial, you’ll build a support-ticket router that assigns a department and priority to each incoming ticket. Along the way, you’ll define a DSPy signature, run a module, create a small labelled dataset, evaluate results, and optimize the application.

This approach is useful when a prompt-based system has become difficult to manage. Rather than repeatedly editing a large prompt by hand, you define the inputs, outputs, and success criteria. DSPy can then search for better instructions and examples for the model you choose.

If you’re completely new to the framework, our existing DSPy tutorial covers the broader foundations. This guide takes a more practical route by building and testing one small business application from end to end.

What DSPy changes in a prompt-based application

A conventional language-model application often starts with a prompt like this:

Read the support ticket. Return the responsible department and the priority. Use one of Billing, Technical Support, Sales, or Account Management. Use low, medium, or high priority.

That can work. The trouble starts when performance is inconsistent.

You might add exceptions for refunds. Then instructions for outages. Then five examples. Then formatting rules. Before long, the prompt becomes a critical business asset that nobody wants to edit because every change can create a regression elsewhere.

DSPy separates the application into clearer parts:

  • Signatures define what goes in and what must come out.
  • Modules define the reasoning pattern, such as direct prediction or step-by-step reasoning.
  • Examples show what good output looks like.
  • Metrics state how you measure success.
  • Optimizers improve instructions and example selection against your metric.
  • Evaluation tests the result against held-out examples.

This does not remove the need for judgement. Someone still needs to define correct outcomes and check that the application is safe to use. It does make the improvement process more structured and repeatable.

For a business owner, the main benefit is control. You can connect changes in model behaviour to a labelled test set and an agreed measure of quality, rather than deciding that a new prompt “seems better.”

When DSPy is worth using

DSPy is most useful when all of these are true:

  1. You have a repeatable language task.
  2. You can describe what a good answer looks like.
  3. You can collect at least a modest set of reviewed examples.
  4. Prompt changes have started taking too much trial and error.
  5. Accuracy matters enough to test before deployment.

Ticket routing, document extraction, lead qualification, policy checks, email categorisation, and internal knowledge-answering are good candidates.

It is less useful for a one-off task. If you need a quick draft of a marketing email, use a straightforward chat prompt. If you have no way to judge the quality of the output, optimization will not solve that underlying problem. DSPy needs a useful definition of success.

For business processes that require operational monitoring after deployment, the design and review workflow behind Omni Ops can help turn a promising prototype into a process people can trust.

Step 1: Set up your project

Create a virtual environment, then install DSPy:

python -m venv .venv
source .venv/bin/activate
pip install dspy

On Windows PowerShell, activate it with:

.venv\Scripts\Activate.ps1

You also need credentials for a language-model provider. Keep the provider key in an environment variable, not in source code or a shared notebook.

Here is a basic configuration pattern:

import os
import dspy

lm = dspy.LM(
    "provider/model-id",
    api_key=os.environ["MODEL_PROVIDER_API_KEY"],
    temperature=0.3,
    max_tokens=250,
)

dspy.configure(lm=lm)

Replace provider/model-id with the identifier required by your chosen provider.

A temperature of 0.3 is a sensible starting point for factual classification and extraction tasks. It allows some flexibility while reducing unnecessary variation. For high-volume routing, use the same model configuration in testing and production. Otherwise, you cannot tell whether a changed result comes from your DSPy program or a changed model setting.

Step 2: Define a signature

A signature is the contract for a task. It describes inputs and outputs without forcing you to write every instruction by hand.

Our support team wants incoming tickets sent to one of four departments:

  • Billing
  • Technical Support
  • Sales
  • Account Management

They also want a priority level of low, medium, or high.

Create a file called ticket_router.py and add this signature:

import dspy

class RouteTicket(dspy.Signature):
    """Route a customer support ticket to the correct department and priority."""

    ticket = dspy.InputField(
        desc="The customer's support request"
    )

    department = dspy.OutputField(
        desc="One of: Billing, Technical Support, Sales, Account Management"
    )

    priority = dspy.OutputField(
        desc="One of: low, medium, high"
    )

The docstring and field descriptions matter. They give DSPy context when it builds instructions for the language model.

Keep output choices constrained. “Choose the best department” sounds reasonable, but it leaves room for outputs such as “Customer Success” or “Finance.” A restricted set gives you cleaner data, easier reporting, and a metric that can score answers automatically.

This is a key difference between an informal prompt and an application contract. Your business process should not need to guess what “urgent-ish” means.

Step 3: Turn the signature into a working module

A module executes a signature. The simplest option is dspy.Predict.

router = dspy.Predict(RouteTicket)

Now call it with a ticket:

result = router(
    ticket=(
        "I was charged twice for my September subscription. "
        "Please refund the duplicate payment."
    )
)

print(result.department)
print(result.priority)

A good result would be:

Billing
medium

At this point, you have a functioning language-model application. DSPy handles the conversion of your signature into a prompt, sends it to the configured model, and structures the response around the named output fields.

Do not assume that the first answer is correct merely because it looks plausible. Classification systems often appear reliable on obvious examples and fail at boundaries. For example, “Our users cannot log in” might be Technical Support and high priority, while “How do I add another user?” is Account Management and low priority. Your test data must include both.

Step 4: Build a small labelled dataset

Optimization only works when it has examples to learn from. Start small, but make the examples representative of actual work.

trainset = [
    dspy.Example(
        ticket="I was charged twice for my monthly subscription.",
        department="Billing",
        priority="medium",
    ).with_inputs("ticket"),

    dspy.Example(
        ticket="Our entire team gets an error whenever we try to sign in.",
        department="Technical Support",
        priority="high",
    ).with_inputs("ticket"),

    dspy.Example(
        ticket="Can you send pricing for 80 users on an annual plan?",
        department="Sales",
        priority="medium",
    ).with_inputs("ticket"),

    dspy.Example(
        ticket="How can I change the person who receives our invoices?",
        department="Account Management",
        priority="low",
    ).with_inputs("ticket"),
]

The .with_inputs("ticket") call is important. It tells DSPy that ticket is supplied to the program, while department and priority are the expected labels.

In a real business application, do not use the same data for training and evaluation. Put aside a separate development set, often called a devset, that the optimizer never sees.

devset = [
    dspy.Example(
        ticket="A payment was taken after we cancelled our account.",
        department="Billing",
        priority="high",
    ).with_inputs("ticket"),

    dspy.Example(
        ticket="The dashboard has been unavailable for two hours.",
        department="Technical Support",
        priority="high",
    ).with_inputs("ticket"),

    dspy.Example(
        ticket="Do you offer a discount for a three-year agreement?",
        department="Sales",
        priority="low",
    ).with_inputs("ticket"),

    dspy.Example(
        ticket="Please update our company name on future invoices.",
        department="Account Management",
        priority="low",
    ).with_inputs("ticket"),
]

Four examples are enough to demonstrate the workflow, not enough to approve a production process. As a practical starting point, collect examples from multiple customer types, wordings, and edge cases. Review them with the people who currently do the work. If your support leaders disagree on the correct department, the model is not the first problem to fix.

Step 5: Write a metric that reflects business value

A metric decides whether a prediction is good.

For this ticket router, an answer is correct only if it gets both department and priority right:

def routing_metric(example, prediction, trace=None):
    correct_department = (
        prediction.department.strip().lower()
        == example.department.strip().lower()
    )

    correct_priority = (
        prediction.priority.strip().lower()
        == example.priority.strip().lower()
    )

    return int(correct_department and correct_priority)

This metric returns 1 for a fully correct answer and 0 otherwise.

That strictness is appropriate if a wrong department or priority causes operational problems. But metrics should reflect your actual workflow.

If the team can tolerate a medium-priority ticket being marked low, but cannot tolerate a security issue being marked low, use a weighted metric. You might give a severe penalty to under-prioritizing critical incidents and a smaller penalty to over-prioritizing routine requests.

Do not optimize for a convenient metric if it conflicts with business risk. A high score on an incomplete metric can hide a poor process.

Step 6: Evaluate the baseline before optimizing

You need a baseline. Run the unoptimized router against the separate development set.

evaluate = dspy.Evaluate(
    devset=devset,
    metric=routing_metric,
    display_progress=True,
)

baseline_score = evaluate(router)
print(f"Baseline score: {baseline_score}")

The score gives you a starting point. More valuable than the number is the error review.

For every incorrect result, capture:

  • The original ticket text
  • The expected department and priority
  • The predicted department and priority
  • Whether the source label was definitely correct
  • Why the model probably made the mistake
  • Whether the issue needs better instructions, better examples, or a process change

This review often finds problems that no optimizer can address. A ticket might need more context than the body text contains. Another may belong to two teams. In those cases, change the workflow. You could add a needs_review output, allow multi-label routing, or send uncertain cases to a human queue.

Step 7: Optimize the router with examples

DSPy optimizers compile a better version of your program against your task and metric. A good first choice for a small, labelled dataset is BootstrapFewShot.

optimizer = dspy.BootstrapFewShot(
    metric=routing_metric
)

optimized_router = optimizer.compile(
    router,
    trainset=trainset
)

The optimizer uses your examples and metric to construct a stronger prompting strategy for the program. It may select useful demonstrations and develop clearer task instructions.

Evaluate the optimized version on the held-out development set:

optimized_score = evaluate(optimized_router)
print(f"Optimized score: {optimized_score}")

The important comparison is baseline versus optimized score on the devset, not on trainset.

If you score on training examples alone, you risk rewarding memorisation. A system that performs well on tickets it has already seen may still route new ticket styles badly.

For more complex programs, DSPy provides other optimization approaches, including MIPROv2. These are worth exploring when you have a better dataset, a reliable metric, and a meaningful performance gap to close. Start with the simplest optimizer that gives you a clear comparison. Complexity without disciplined evaluation only makes failures harder to diagnose.

Step 8: Add a review path before production

A routing model should not silently make every decision. Give the business a way to see what happened and correct it.

You can extend the signature with a confidence-style output and a short rationale for internal review:

class RouteTicketWithReview(dspy.Signature):
    """Route a customer support ticket and flag uncertain cases."""

    ticket = dspy.InputField(desc="The customer's support request")

    department = dspy.OutputField(
        desc="One of: Billing, Technical Support, Sales, Account Management"
    )

    priority = dspy.OutputField(
        desc="One of: low, medium, high"
    )

    needs_review = dspy.OutputField(
        desc="yes if the ticket is ambiguous or lacks enough information, otherwise no"
    )

    reason = dspy.OutputField(
        desc="A short explanation for an internal support reviewer"
    )

Do not show an internal explanation to customers without reviewing the design. It may be inaccurate, expose internal logic, or cause the team to place too much trust in a fluent answer.

For a production workflow, log the input, result, version of the program, model configuration, final human action, and any correction. Those corrections become the next source of labelled examples.

This is where a prototype becomes an operating system for a business process. Our Omni work focuses on that broader practical question: how to connect language-model capabilities to real workflows, ownership, and measurable outcomes.

Common DSPy mistakes and how to avoid them

Optimizing without enough good labels

DSPy cannot repair weak ground truth. If your examples are inconsistent, the optimizer learns inconsistency.

Create clear routing definitions first. Have two experienced staff members independently label a sample. Discuss disagreements and document the rule. Then build the dataset.

Using vague output fields

“Return the right team” is vague. “Return exactly one of Billing, Technical Support, Sales, Account Management” is testable.

Constrained outputs reduce cleaning work and make it possible to calculate a meaningful score.

Testing on the same examples used for optimization

This is the most common evaluation error. Keep a development set separate from the training set. For higher-stakes applications, maintain a third untouched test set for final approval.

Treating one metric as the whole truth

Exact-match accuracy is useful, but it may not represent business cost. Review confusion patterns. Routing Sales inquiries to Account Management may be inconvenient. Routing a widespread outage to Sales could be much more serious.

Changing several things at once

Do not change the model, temperature, signature, training data, and metric in one experiment. You will not know what caused the result.

Change one meaningful variable, rerun evaluation, and keep a simple experiment log.

Forgetting that source data can change

Support topics shift. Product releases create new terminology. Policy changes alter what counts as urgent.

Re-evaluate on recent examples at a regular cadence. This is particularly important after a major process change or when staff report an increase in incorrect routing.

Where to go after your first DSPy application

Once the ticket router is working, extend it carefully. You could add a module that extracts account identifiers, a retrieval step that finds relevant support policies, or a second check for high-risk requests. Evaluate each addition independently before combining them.

You can find more practical implementation topics in our guides library, while Enterprise DNA Learn offers structured material for teams building stronger data and automation capability.

DSPy improves prompt-based systems when you can define success, collect reviewed examples, and test changes against real work. It is not a substitute for process design. It is a better way to turn a process into a language-model application that can be measured and improved.

If you’re deciding where a DSPy-style application fits in your business, book a call with Sam to discuss the workflow, data, and evaluation plan before you build.