Book a consultation

The incident runbook for production AI

A practical incident process for detecting, containing, diagnosing, correcting, and learning from failures in agents and RAG systems.

The incident runbook for production AI

At 10:14, customer feedback on a banking assistant drops.

At 10:19, the on-call engineer filters the affected traces and sees every bad answer citing the same policy version.

At 10:24, the team disables automated answers for that intent and routes requests to staff.

At 10:41, they confirm that the updated policy never reached the vector index.

This is what an AI incident process should make possible. The exact times are illustrative, but the sequence is not. Detect, contain, diagnose, correct, verify, and learn.

A working production incident process connects evaluation dashboards, traces, prompt versions, human handoff, and a growing test library. That connection is more useful than a separate "AI monitoring" dashboard nobody has integrated with on-call work.

AI incident response loop from detection through containment, diagnosis, correction, and learning
Every incident should improve the next release.
Circular quality feedback loop from build through evaluation, shipping, tracing, and learning
Incident response feeds evaluation and changes the next release.

Define what counts as an AI incident

Not every bad answer is an incident. Not every incident returns an error.

Create incident classes before launch:

quality:
  examples: [wrong_policy, unsupported_claim, systematic_misroute]

safety:
  examples: [pii_disclosure, unauthorized_action, prompt_injection_effect]

data:
  examples: [stale_index, missing_source, permission_filter_failure]

operations:
  examples: [provider_outage, tool_timeout, runaway_cost, checkpoint_failure]

business:
  examples: [conversion_drop, false_deflection, excessive_human_escalation]

Define severity using impact, exposure, reversibility, and duration. A single unauthorized refund can be more severe than 1,000 awkward sentences.

Assign an incident owner and domain owner for each class. The on-call engineer may contain the system while a policy owner decides which answer is correct.

Detect quality regressions with several signals

Infrastructure metrics alone will miss most semantic incidents.

Use four signal groups:

  • Service signals include availability, latency, errors, rate limits, and queue depth.

  • Behavior signals include tool-call count, loops, retries, fallbacks, and approval bypass.

  • Quality signals include evaluator scores, groundedness, refusal accuracy, and escalation.

  • Business signals include customer feedback, resolution, correction rate, and cost per successful task.

No single signal is reliable enough. A fall in customer satisfaction may follow an unpopular but correct policy. A groundedness judge may miss that retrieval supplied an outdated source. Correlate signals through trace cohorts.

Alert on hard safety events immediately. Use windows and baselines for noisy quality metrics. Attach the affected release, intent, tenant, and representative traces to the alert.

Containment must be designed before the incident

The fastest safe action depends on the failure class.

bad prompt version       roll back behavior bundle
stale source             disable intent or pin previous valid source
tool outage              open circuit and use read-only fallback
unauthorized write       disable capability and revoke credential
judge degradation        stop automated promotion, keep service running
provider outage          route to tested fallback or human queue
runaway cost             enforce budget and terminate long runs

Containment controls need to exist in code and operations:

  • kill switch per workflow and capability
  • feature flags per tenant or intent
  • model and prompt rollback
  • circuit breakers for tools and providers
  • human-routing fallback
  • read-only mode
  • queue pause and replay policy

A global shutdown is sometimes necessary, but narrow controls reduce customer impact. If only refund execution is unsafe, the assistant may continue reading policy and drafting responses.

Test kill switches during normal operation. A control first used during an incident is only a theory.

Diagnose from a saved trace cohort

One trace can mislead. Build a cohort of affected and unaffected runs.

Filter by the time window, intent, release bundle, source version, model, tool result, and evaluator label. Compare the first failing cohort with the last healthy one.

healthy traces
  workflow quote-agent-2.8.0
  policy index 30
  groundedness 0.97

failing traces
  workflow quote-agent-2.8.0
  policy index 30
  source registry says index 31 should be active

The unchanged workflow and stale index narrow the search. The model does not need to be the first suspect.

LangSmith traces contain nested runs, metadata, feedback, and tags. Tag confirmed incidents and save the trace query used for diagnosis. This preserves the evidence behind the root-cause decision.

Review observable behavior, not private chain-of-thought. Inputs, retrieved documents, tool calls, state transitions, versions, validation, and output are enough to reconstruct most failures.

Use a hypothesis table

Incident debugging gets slower when everyone changes prompts at once.

hypothesis            evidence needed                  owner
stale retrieval       source and index versions        data platform
model regression      same cases on prior model        AI platform
tool failure          attempt spans and status         service owner
permission defect     policy decision logs             security
prompt regression     pairwise run on prior prompt     agent team

Test one hypothesis against the saved cohort. Do not deploy speculative fixes to production traffic.

If the incident involves a side effect, reconcile downstream state before retrying. A timeout does not prove an operation failed.

Correct the smallest responsible component

Fix the source of the failure, then add defense in depth where useful.

For stale retrieval, restore the missing index update and verify propagation. A prompt that tells the model to be careful about dates is not the root fix.

For a repeated tool call, correct loop or retry logic and add idempotency. Asking the model to "avoid duplicate calls" is weaker than enforcing an operation budget.

For an unauthorized action, remove the capability or fix downstream authorization. A refusal example in the prompt is not a permission boundary.

Record every changed artifact in the candidate behavior bundle. The incident fix may include data, code, prompt, graph, evaluator, or policy changes.

Verify with a regression case and a broader suite

Turn the confirmed incident into a versioned dataset example.

The example should include the input, expected output or action, source versions, permission profile, expected trajectory, and failure label. Add nearby variants so the system does not memorize one sentence.

Run three checks:

  1. The candidate passes the exact incident case.
  2. The candidate passes related failure-class cases.
  3. The full suite shows no unacceptable regression.

LangSmith's evaluation workflow supports moving failing production traces into a dataset, validating a fix offline, and redeploying. Keep the original trace linked to the regression example so future reviewers know why the case exists.

Canary the fix on the affected intent. Compare the candidate with the healthy baseline. Do not close the incident when the code merges.

Integrate with the existing incident system

AI incidents need the same paging, ownership, escalation, communication, and postmortem process as other production incidents.

Send alerts to the company's incident-management or ITSM system. Include:

severity and incident class
affected workflow and release
time window and estimated exposure
containment status
trace cohort link
business owner
technical owner
rollback target

Do not create a parallel AI-only process that security and operations cannot see.

The runbook should name the person authorized to disable a model route, revoke a tool credential, or force human handling. Those decisions are too slow when ownership is discovered during the incident.

Preserve evidence without preserving everything forever

An incident may require longer retention for selected traces. Apply a legal or security hold to the relevant cohort rather than extending retention for all prompts and documents.

Store hashes or references where raw data is unnecessary. Restrict access to incident evidence. Record who exported or reviewed sensitive traces.

If a customer requests deletion, the incident process still needs a documented rule for what can be retained under legal obligations. That is a policy decision to make with counsel, not during an outage.

Close with a change to the system

A useful postmortem produces concrete work:

  • a regression example
  • a new or corrected evaluator
  • a missing trace attribute
  • a tighter permission
  • a tested containment control
  • an ownership or runbook correction

Avoid "the model hallucinated" as a root cause. It describes an output, not why the system allowed that output to reach a user or action.

The incident is closed when the fix is deployed, monitoring confirms recovery, affected downstream state is reconciled, and prevention work has owners.

The final article in this series moves these controls to the beginning of delivery: Select the model last

Start with the work

Bring the workflow that needs attention.

Pick a time for a working session or send a short brief. Either way, we will come prepared to understand where the work gets stuck.

Talk through the work

Book a 20-minute consultation.

Bring the workflow that feels slow or fragile. We will determine whether it is a sensible candidate for an AI system.

Bartosz LuderaBartosz LuderaFounder, Harnessloop

Choose a time for a 20-minute consultation.

Send a workflow brief

Prefer to write it down?

Tell us where work waits, repeats, or falls through the cracks.