Evaluation is the specification: how to test AI systems before users do
A production evaluation architecture for AI agents, including datasets, deterministic checks, LLM judges, trajectory tests, and release gates.

release_gate:
eligible_case_deflection: ">= 60%"
critical_policy_errors: "0"
tool_argument_validity: ">= 99.5%"
p95_latency: "< 4.5s"
cost_per_resolved_case: "< $0.18"
This is a better starting point for an AI project than a model shortlist.
It forces the team to answer the question that most prototypes postpone. What must the system do well enough for us to release it, and which failures are unacceptable even if the average score looks good?
Evaluation is the closest thing an AI system has to an executable specification. A conventional specification says which outputs follow from known inputs. An AI specification also has to describe distributions, tolerances, abstention, and the behavior of the execution path.
An agent can produce a correct final sentence after calling the wrong customer database three times. Output-only evaluation calls that a success. Production engineering cannot.
Start with the loss, not the score
"Accuracy above 90%" is not a production contract until you know which 10% fails.
Suppose a bank assistant handles 1,000 requests. It answers 920 correctly, escalates 50 uncertain cases, gives 25 harmless incomplete answers, and exposes private account data in five responses. Its accuracy is 92%. The system still cannot ship.
A useful evaluation design starts with a failure taxonomy and the cost of each class.
from enum import StrEnum
class FailureClass(StrEnum):
WRONG_POLICY = "wrong_policy"
UNAUTHORIZED_DISCLOSURE = "unauthorized_disclosure"
INVALID_TOOL_ARGUMENTS = "invalid_tool_arguments"
MISSED_ESCALATION = "missed_escalation"
DUPLICATE_SIDE_EFFECT = "duplicate_side_effect"
POOR_STYLE = "poor_style"
The order is intentional. A style defect and an unauthorized disclosure do not belong in the same average.
I prefer two kinds of release gates:
-
Hard constraints block a deployment when a critical failure appears. Privacy exposure, an unauthorized write, or a duplicate financial action usually belongs here.
-
Optimization metrics compare versions after both have passed the hard constraints. Task completion, latency, cost, and escalation rate fit this group.
This prevents a common mistake where an improvement in a high-volume, low-risk metric compensates for a rare but severe regression.
Define the eligible population
Deflection and task-completion metrics are easy to inflate by changing the denominator.
If an assistant resolves 600 of 1,000 requests, the apparent deflection rate is 60%. But perhaps only 700 requests were eligible for automation and 300 required a regulated human decision. The agent resolved 600 of 700 eligible requests, which is 85.7%. It also needs a separate metric for correctly identifying the 300 ineligible cases.
Your evaluation schema should record that distinction.
{
"case_id": "support-1842",
"intent": "overdraft_fee_dispute",
"automation_eligible": false,
"required_action": "escalate",
"required_tools": ["get_account_status", "get_fee_policy"],
"forbidden_tools": ["waive_fee"],
"risk_tier": "high",
"source_version": "fee-policy-2026-08-14"
}
This row specifies more than an ideal answer. It describes the allowed trajectory and the operating boundary.
Build datasets around decisions people make
The first dataset should come from real work, not prompts invented during development.
Collect resolved tickets, approved documents, completed analyses, and the decisions made by experienced operators. Ask domain experts why they escalated a case and which facts changed the answer. Their corrections often contain more useful evaluation data than the final prose.
Coverage needs several axes:
- intent and task type
- frequency in production
- risk tier
- input quality
- data source and permission profile
- expected tool path
- language and locale when relevant
Do not let the dataset mirror production frequency exactly. A rare privacy failure may occur in 0.01% of traffic, but it deserves far more than 0.01% of a release suite. Keep a representative set for estimating normal performance and a challenge set weighted toward dangerous cases.
Version the dataset. A score without a dataset version is not reproducible. LangSmith datasets are versioned, which lets an experiment retain the exact examples used at run time rather than silently inheriting later edits.
Use three evaluation layers
A production evaluation stack needs deterministic, semantic, and behavioral layers. Each catches a different kind of failure.
Layer 1 uses deterministic checks
Code should evaluate anything code can know with certainty.
def validate_quote(run: dict, reference: dict) -> dict:
output = run["outputs"]
passed = (
output["currency"] == reference["currency"]
and output["total"] == reference["total"]
and output["customer_id"] == reference["customer_id"]
and set(output["citations"]) >= set(reference["required_sources"])
)
return {"key": "quote_contract", "score": int(passed)}
Schema validation, permissions, exact identifiers, numeric reconciliation, tool allowlists, and rate limits are cheaper and more repeatable than an LLM judge. If a deterministic check can settle the question, adding a judge introduces cost and disagreement without adding information.
Layer 2 evaluates semantics
Some outputs can be correct in more than one form. An LLM judge can compare an answer with source material or score it against a rubric.
The rubric needs operational definitions. "Is this answer good?" is not a useful judge prompt. A groundedness rubric should tell the judge to identify each externally verifiable claim, find support in the supplied context, and fail the answer if any material claim lacks support.
Run judge calibration before trusting its scores:
- Have domain experts label a sample without seeing the judge output.
- Measure agreement per failure class, not only overall agreement.
- Inspect false passes on high-risk cases.
- Add counterexamples where the judge was confidently wrong.
- Recalibrate after changing the judge model or rubric.
An 85% agreement rate may be acceptable for routing low-risk cases to review. It is not enough if most disagreement occurs on policy violations.
LangSmith supports code, human, LLM-as-judge, and pairwise evaluators. Pairwise evaluation is often easier to calibrate when the question is whether candidate B is better than the current production version. It still needs order randomization because judges can prefer the first or second answer.
Layer 3 evaluates behavior
Agent evaluation has to inspect the path, not only the destination.
Useful trajectory checks include:
- whether the agent selected the right tool
- whether arguments were valid and authorized
- whether it repeated a call without new information
- whether it exceeded the action budget
- whether a risky action passed through approval
- whether it stopped after reaching a terminal state
A balance lookup that hits the database three times may return the right number. Under real traffic, those calls increase latency and database load. If one call has a side effect, repetition becomes a correctness defect.
Behavioral evaluators should consume trace data. LangSmith experiments retain execution traces alongside outputs and scores, so a failed trajectory can be inspected without recreating the run from application logs.
Separate offline evidence from online signals
Offline and online evaluation answer different questions.
Offline evaluation asks whether a candidate version passes known cases. It has reference outputs, expected tools, and labeled failure classes. Run it on prompt changes, model changes, tool-schema updates, retrieval changes, and graph edits.
Online evaluation asks whether production traffic is moving away from what you tested. Most production requests have no reference answer, so online evaluators rely on signals such as policy checks, groundedness against retrieved context, unusual trajectory length, negative feedback, and changes in escalation rate.
The two systems need a promotion path:
online signal
↓
trace review
↓
confirmed failure class
↓
versioned dataset example
↓
offline regression test
↓
candidate fix
LangSmith documents this exact feedback loop. Production traces reveal cases the curated set missed. Those cases become offline tests, and the offline suite verifies the fix before release.
Do not add every negative user reaction directly to the dataset. A thumbs-down can mean the answer was wrong, the policy was unpopular, or the user disliked being denied. Review and label the trace first.
Keep evaluation affordable without making it decorative
A 500-row dataset can trigger thousands of paid model and judge calls for one pull request. Teams then sample so aggressively that the suite stops protecting the release.
Use evaluation tiers:
pull request
deterministic suite
risk-focused semantic sample
changed-component trajectory tests
merge to main
full regression suite
pairwise comparison with production
scheduled
broad judge suite
adversarial and long-tail cases
Select pull-request cases by dependency and failure class. A retriever change should run retrieval and groundedness cases. A tool-schema change should run argument, permission, and trajectory cases. Always include a small fixed sentinel set so unrelated regressions remain visible.
Cache model outputs only when the evaluated application version and inputs are identical. Do not reuse a cached judge score after changing the rubric or judge model.
Report cost per experiment and cost per caught regression. Evaluation spend is easier to defend when the team can connect it to prevented incidents.
Read the score as a distribution
The difference between 92.0% and 93.0% on 100 examples may be noise. Ten changed answers can move the metric by ten percentage points.
For every comparison, keep the paired result per example. Count regressions, improvements, and unchanged cases. Add confidence intervals for aggregate rates. Break results down by risk tier and intent.
overall task success 93.1% (+1.2 pp)
high-risk task success 88.0% (-4.0 pp)
critical failures 1 (+1)
median tool calls 2 (-1)
cost per success $0.14 (-$0.03)
The candidate above should not ship despite the better overall score. The high-risk slice regressed and a critical failure appeared.
The release decision should be boring
By the time a model reaches staging, the team should know the dataset, evaluator versions, hard constraints, and acceptable tradeoffs. The release review becomes a comparison against a contract, not a live debate about whether the answers "feel better."
That is the practical value of treating evaluation as specification. It moves quality decisions before implementation and leaves an evidence trail after deployment.
The next article in this series covers the data that makes evaluation and debugging possible: Your agent returned the wrong answer. Can you explain why?

