Prompts and models are production dependencies
How to version, evaluate, canary, roll back, and audit prompt and model changes without turning production into the test environment.

system_prompt_final_v7_really_final.txt
The joke is old because the failure mode is common.
A prompt changes. The assistant sounds better in three manual tests. Somebody deploys it. Two days later, tool-call volume rises, escalation drops for the wrong reason, and one policy case starts failing.
No code changed, so the release process did not treat it as a release.
Prompts need the same change discipline as code, and provider benchmarks cannot select a model for your private workload. Treat the application binary and its behavior as one release.
Define the behavior bundle
An agent version is a bundle of coupled artifacts.
agent_release: quote-agent-2.8.1
code_commit: 4ca128f
graph_version: quote-flow-12
system_prompt: quote-system-17
tool_schema: quote-tools-9
retrieval_config: proposals-31
model_route: quote-router-6
policy_bundle: commercial-policy-22
evaluator_suite: quote-evals-14
dataset_version: quote-gold-2026-09-02
Any of these can change output or trajectory. Record the full bundle in traces and experiment metadata.
Versioning only the prompt text misses template variables, tool descriptions, model parameters, middleware, retrieval filters, and graph topology. A reproducible release needs the rendered configuration and references to immutable artifacts.
Store prompts as reviewed source
Prompts need an owner, tests, and a reason for change.
Keep templates in source control or a prompt registry with immutable versions. Do not let the production application load an unversioned "latest" prompt. Production should pin a version and move the pointer through deployment.
A useful change record explains the observed failure and expected effect:
change: quote-system-16-to-17
reason: "Agent omitted warranty exclusions for refurbished products"
evidence:
- trace: 8f72
- dataset_cases: [warranty-44, warranty-79]
expected_effect:
- "Include exclusions when product condition is refurbished"
non_goals:
- "No change to pricing or discount behavior"
owner: sales-platform
"Improve the prompt" is not a reviewable reason.
Keep secrets and environment-specific values outside prompt templates. A prompt registry is not a secrets manager. If the model needs a value for a tool call, pass a reference through a typed interface rather than interpolating credentials into context.
Test the change at three scopes
A prompt edit should run component, workflow, and production-shaped tests.
Component tests check whether required instructions render and structured outputs validate.
Workflow tests run the agent against cases affected by the change and inspect tool trajectories.
Production-shaped tests run the broader regression dataset with realistic retrieval, permissions, latency, and tool behavior.
prompt unit checks
↓
targeted failure-class evals
↓
full offline experiment
↓
staging with real integrations
↓
shadow or canary traffic
Use pairwise evaluation against the current production release as well as absolute gates. A candidate may remain above the minimum threshold while still getting worse on a high-volume intent.
LangSmith experiments retain outputs, evaluator scores, and traces for each dataset example. Add the behavior bundle to experiment metadata so a comparison identifies exactly which prompt, model, graph, and index produced each result.
Select models on your cases
Public benchmarks answer questions about public tasks under someone else's harness. They do not know your tool schemas, abbreviations, policy documents, latency region, or cost model.
Build a model bake-off on the same versioned dataset and application configuration.
candidate A
task success 94.2%
critical failures 0
p95 latency 3.8s
cost per success $0.21
candidate B
task success 93.7%
critical failures 0
p95 latency 2.1s
cost per success $0.08
Candidate B may be the better production choice if the 0.5-point difference falls within uncertainty and latency changes user behavior. The decision belongs to the production contract, not a leaderboard rank.
Run multiple repetitions on cases where model variance matters. Keep temperature and other parameters explicit. Test rate limits and provider failures as well as output quality.
Treat a provider upgrade as untrusted
A provider can update an aliased model without changing your code. Pin dated versions where available. When pinning is impossible, monitor behavior for drift and keep a tested fallback.
Model change management should cover:
- announced deprecations
- silent behavior changes behind aliases
- pricing changes
- regional availability
- latency and rate-limit changes
- safety-policy changes
- structured-output and tool-calling differences
Do not assume the fallback is compatible because it accepts the same API format. Run the same evaluation suite against it. Verify tool arguments and stop behavior, not only final answers.
A fallback can create a compliance problem if it runs in another region or retains data differently. Routing policy needs legal and security metadata alongside quality and cost.
Make structured outputs part of the contract
Free-form text creates a weak boundary between model and application.
Use schemas for decisions and tool proposals:
class QuoteDecision(BaseModel):
outcome: Literal["approve", "review", "reject"]
reasons: list[str]
policy_ids: list[str]
confidence: float
LangChain structured output can use provider-native schemas where supported or tool calling where it is not. The application receives validated data rather than parsing prose.
Schema validation does not prove the decision is correct. It makes failures visible and keeps invalid output away from downstream code. Track validation retries because a candidate model that needs three attempts is slower and more expensive than its final output suggests.
Version schemas separately. A prompt compatible with QuoteDecisionV2 may break a consumer expecting V1 even when every field is valid.
Canary a behavior bundle
A canary should isolate the candidate bundle and label every trace.
Start with internal or shadow traffic. Then route a small percentage of eligible low-risk cases. Keep users sticky to one version during a conversation so context does not cross behavior bundles.
Compare:
- hard safety constraints
- task success and escalation
- tool trajectory
- latency and cost
- user correction and override
- downstream business outcome where available
Stop automatically on critical failures. Do not wait for statistical significance after a privacy leak or unauthorized write.
Rollback should move the release pointer to the previous complete bundle. Rolling back the prompt while leaving the new tool schema or index active may not restore old behavior.
Version judges and datasets too
The measurement system changes just like the application.
A stricter judge can make production quality appear to drop even if the agent did not change. A corrected dataset label can move a score. Record judge model, rubric, examples, and dataset version with every experiment.
When a judge changes, run it against a human-labeled calibration set and compare the old and new disagreements. Do not rewrite historical experiment scores without preserving which evaluator produced them.
This is one reason to keep raw trace references and immutable experiment results. The team can re-evaluate old runs with a new judge while retaining the original score.
Keep the main branch protected by evidence
Full behavioral suites can be expensive. Use tiered execution without making pull-request checks meaningless.
pull request
syntax and schema checks
deterministic evaluators
impacted failure classes
small fixed sentinel set
main branch
full offline suite
pairwise production comparison
security and trajectory suite
scheduled
long-tail and adversarial suite
fallback model verification
Map artifacts to affected tests. A policy prompt change should run policy cases. A model route change should run the whole suite because its effect is broad.
Require the change record, experiment link, and rollback target in the pull request. The reviewer should be able to see which cases improved and regressed, not only the aggregate score.
A prompt deployment is a software deployment
The implementation details differ, but the operating standard should be familiar. Review the change, run tests, deploy to staging, canary it, watch traces, and keep a rollback path.
The useful extension is that the deployed artifact includes data, model, prompt, tools, and graph. That bundle is what users experience.
The next article covers what happens when a tested bundle still fails: The incident runbook for production AI

