The incident runbook for production AI
A practical incident process for detecting, containing, diagnosing, correcting, and learning from failures in agents and RAG systems.

At 10:14, customer feedback on a banking assistant drops.
At 10:19, the on-call engineer filters the affected traces and sees every bad answer citing the same policy version.
At 10:24, the team disables automated answers for that intent and routes requests to staff.
At 10:41, they confirm that the updated policy never reached the vector index.
This is what an AI incident process should make possible. The exact times are illustrative, but the sequence is not. Detect, contain, diagnose, correct, verify, and learn.
A working production incident process connects evaluation dashboards, traces, prompt versions, human handoff, and a growing test library. That connection is more useful than a separate "AI monitoring" dashboard nobody has integrated with on-call work.
Define what counts as an AI incident
Not every bad answer is an incident. Not every incident returns an error.
Create incident classes before launch:
quality:
examples: [wrong_policy, unsupported_claim, systematic_misroute]
safety:
examples: [pii_disclosure, unauthorized_action, prompt_injection_effect]
data:
examples: [stale_index, missing_source, permission_filter_failure]
operations:
examples: [provider_outage, tool_timeout, runaway_cost, checkpoint_failure]
business:
examples: [conversion_drop, false_deflection, excessive_human_escalation]
Define severity using impact, exposure, reversibility, and duration. A single unauthorized refund can be more severe than 1,000 awkward sentences.
Assign an incident owner and domain owner for each class. The on-call engineer may contain the system while a policy owner decides which answer is correct.
Detect quality regressions with several signals
Infrastructure metrics alone will miss most semantic incidents.
Use four signal groups:
-
Service signals include availability, latency, errors, rate limits, and queue depth.
-
Behavior signals include tool-call count, loops, retries, fallbacks, and approval bypass.
-
Quality signals include evaluator scores, groundedness, refusal accuracy, and escalation.
-
Business signals include customer feedback, resolution, correction rate, and cost per successful task.
No single signal is reliable enough. A fall in customer satisfaction may follow an unpopular but correct policy. A groundedness judge may miss that retrieval supplied an outdated source. Correlate signals through trace cohorts.
Alert on hard safety events immediately. Use windows and baselines for noisy quality metrics. Attach the affected release, intent, tenant, and representative traces to the alert.
Containment must be designed before the incident
The fastest safe action depends on the failure class.
bad prompt version roll back behavior bundle
stale source disable intent or pin previous valid source
tool outage open circuit and use read-only fallback
unauthorized write disable capability and revoke credential
judge degradation stop automated promotion, keep service running
provider outage route to tested fallback or human queue
runaway cost enforce budget and terminate long runs
Containment controls need to exist in code and operations:
- kill switch per workflow and capability
- feature flags per tenant or intent
- model and prompt rollback
- circuit breakers for tools and providers
- human-routing fallback
- read-only mode
- queue pause and replay policy
A global shutdown is sometimes necessary, but narrow controls reduce customer impact. If only refund execution is unsafe, the assistant may continue reading policy and drafting responses.
Test kill switches during normal operation. A control first used during an incident is only a theory.
Diagnose from a saved trace cohort
One trace can mislead. Build a cohort of affected and unaffected runs.
Filter by the time window, intent, release bundle, source version, model, tool result, and evaluator label. Compare the first failing cohort with the last healthy one.
healthy traces
workflow quote-agent-2.8.0
policy index 30
groundedness 0.97
failing traces
workflow quote-agent-2.8.0
policy index 30
source registry says index 31 should be active
The unchanged workflow and stale index narrow the search. The model does not need to be the first suspect.
LangSmith traces contain nested runs, metadata, feedback, and tags. Tag confirmed incidents and save the trace query used for diagnosis. This preserves the evidence behind the root-cause decision.
Review observable behavior, not private chain-of-thought. Inputs, retrieved documents, tool calls, state transitions, versions, validation, and output are enough to reconstruct most failures.
Use a hypothesis table
Incident debugging gets slower when everyone changes prompts at once.
hypothesis evidence needed owner
stale retrieval source and index versions data platform
model regression same cases on prior model AI platform
tool failure attempt spans and status service owner
permission defect policy decision logs security
prompt regression pairwise run on prior prompt agent team
Test one hypothesis against the saved cohort. Do not deploy speculative fixes to production traffic.
If the incident involves a side effect, reconcile downstream state before retrying. A timeout does not prove an operation failed.
Correct the smallest responsible component
Fix the source of the failure, then add defense in depth where useful.
For stale retrieval, restore the missing index update and verify propagation. A prompt that tells the model to be careful about dates is not the root fix.
For a repeated tool call, correct loop or retry logic and add idempotency. Asking the model to "avoid duplicate calls" is weaker than enforcing an operation budget.
For an unauthorized action, remove the capability or fix downstream authorization. A refusal example in the prompt is not a permission boundary.
Record every changed artifact in the candidate behavior bundle. The incident fix may include data, code, prompt, graph, evaluator, or policy changes.
Verify with a regression case and a broader suite
Turn the confirmed incident into a versioned dataset example.
The example should include the input, expected output or action, source versions, permission profile, expected trajectory, and failure label. Add nearby variants so the system does not memorize one sentence.
Run three checks:
- The candidate passes the exact incident case.
- The candidate passes related failure-class cases.
- The full suite shows no unacceptable regression.
LangSmith's evaluation workflow supports moving failing production traces into a dataset, validating a fix offline, and redeploying. Keep the original trace linked to the regression example so future reviewers know why the case exists.
Canary the fix on the affected intent. Compare the candidate with the healthy baseline. Do not close the incident when the code merges.
Integrate with the existing incident system
AI incidents need the same paging, ownership, escalation, communication, and postmortem process as other production incidents.
Send alerts to the company's incident-management or ITSM system. Include:
severity and incident class
affected workflow and release
time window and estimated exposure
containment status
trace cohort link
business owner
technical owner
rollback target
Do not create a parallel AI-only process that security and operations cannot see.
The runbook should name the person authorized to disable a model route, revoke a tool credential, or force human handling. Those decisions are too slow when ownership is discovered during the incident.
Preserve evidence without preserving everything forever
An incident may require longer retention for selected traces. Apply a legal or security hold to the relevant cohort rather than extending retention for all prompts and documents.
Store hashes or references where raw data is unnecessary. Restrict access to incident evidence. Record who exported or reviewed sensitive traces.
If a customer requests deletion, the incident process still needs a documented rule for what can be retained under legal obligations. That is a policy decision to make with counsel, not during an outage.
Close with a change to the system
A useful postmortem produces concrete work:
- a regression example
- a new or corrected evaluator
- a missing trace attribute
- a tighter permission
- a tested containment control
- an ownership or runbook correction
Avoid "the model hallucinated" as a root cause. It describes an output, not why the system allowed that output to reach a user or action.
The incident is closed when the fix is deployed, monitoring confirms recovery, affected downstream state is reconciled, and prevention work has owners.
The final article in this series moves these controls to the beginning of delivery: Select the model last

