Select the model last: an evaluation-first AI delivery plan
An eight-week delivery plan that defines quality, data, traces, tool boundaries, and operational ownership before selecting the production model.

Most AI delivery plans place model selection in the first week. Reverse that order.
Define the production contract first. Build the evaluation set, tracing, data pipeline, and tool boundaries next. Compare models only when the team can measure them against the work the system will perform.
The calendar in this article uses eight weeks as a planning device. Model selection becomes a short experiment near the end because the measurement system already exists.
Week zero decides whether the project deserves eight weeks
Before staffing delivery, write the operating hypothesis.
For eligible retail-support questions,
the assistant will resolve at least 60 percent without human handling,
while preserving the current critical-error rate,
and keeping cost below the existing cost per resolved contact.
Define the eligible population. Exclude tasks that require regulated advice, unusual discretion, or unavailable data. Estimate volume, current handling time, transfer rate, and cost.
Set a kill condition. If the required data cannot be accessed, critical errors remain above the threshold, or unit economics fail under realistic load, stop. A delivery process without a kill condition turns evidence into a justification for more work.
Assign a business owner for the outcome and a technical owner for the system. The AI team cannot own customer-support policy on behalf of the support organization.
Weeks one and two build the specification
Collect real cases from the people doing the work.
There is no magic dataset size. Start with enough coverage to represent common intents, high-risk cases, missing data, ambiguous wording, and escalation. Fifty carefully reviewed examples can expose more than 500 synthetic variations of the same happy path.
Each example should contain more than a question and ideal answer:
{
"intent": "overdraft_fee_dispute",
"eligible_for_automation": false,
"expected_outcome": "escalate",
"required_sources": ["fee-policy-current"],
"required_tools": ["get_account_status"],
"forbidden_actions": ["waive_fee"],
"risk_tier": "high"
}
Define deterministic, semantic, and behavioral evaluators. Calibrate judge-based metrics with domain experts. Set hard constraints for privacy, authorization, and critical policy errors. Set optimization targets for completion, latency, escalation, and cost.
Create a baseline with the existing human process where possible. The AI system is not competing with an imaginary perfect operator. It is changing a real queue with real error and transfer rates.
LangSmith offline evaluation can store the dataset, run application versions as experiments, and retain scores with execution traces. The tool is useful because the artifact is not a spreadsheet of final answers. It is a repeatable harness.
Weeks three and four make data and execution visible
Build the source registry and ingestion lifecycle before tuning retrieval.
For every source, record ownership, permissions, version, freshness target, and deletion behavior. Make index activation conditional on successful processing and verification. Add a trace link from retrieved chunks to source and ingestion versions.
Instrument the full request path:
identity
↓
intent and routing
↓
retrieval and tools
↓
model calls
↓
validation
↓
approval or response
Record workflow, prompt, model, tool-schema, index, and policy versions. Redact sensitive content before trace export.
At this stage, use a convenient model to exercise the system without optimizing around its quirks. This phase is complete when a wrong answer can be traced to a component and reproduced against the dataset.
LangSmith tracing provides agent-specific runs and feedback. OpenTelemetry can carry traces across the rest of the service stack. A team may use either or both, but it needs stable trace and operation IDs across model, application, and downstream tools.
Weeks five and six define control and failure
Turn tool access into explicit contracts.
List what the agent can read, propose, and execute. Add downstream authorization. Make write operations idempotent. Define retry, timeout, and compensation behavior per tool.
Build the smallest orchestration graph that completes the task. Keep known rules deterministic. Persist state outside conversation history.
classify request
↓
eligible? ── no ──→ human queue
↓ yes
retrieve policy and account data
↓
draft answer
↓
validate policy and privacy
↓
confidence below threshold? ── yes ──→ human queue
↓ no
respond
Use LangGraph persistence when workflows need checkpoints, resume, or human review. Use LangChain middleware for call limits, retries, fallbacks, PII controls, and approval policies where they fit.
Write the incident matrix now:
model unavailable tested fallback or human queue
retrieval empty abstain and escalate
source stale disable affected intent
tool timeout bounded retry and reconciliation
quality drops freeze rollout and inspect traces
unauthorized action disable tool and revoke credential
Run a game day. Break one dependency and verify that the workflow fails as designed.
Week seven runs the model bake-off
Now model selection has an input, harness, and decision rule.
Run each candidate against the same dataset and behavior bundle. Compare hard constraints first. Then compare task success, trajectory, latency, cost, and variance.
candidate A
eligible resolution 63%
critical errors 0
p95 latency 4.2s
cost per resolution $0.17
candidate B
eligible resolution 65%
critical errors 1
p95 latency 3.1s
cost per resolution $0.11
Candidate A wins if one critical error violates the production contract. Aggregate quality cannot compensate for a hard safety failure.
Run pairwise comparisons on changed answers and have domain experts inspect disagreements. Repeat non-deterministic cases. Load-test provider limits and measure queue behavior under fallback.
Keep business logic outside model-specific prompt tricks. A framework abstraction such as LangChain's model interface reduces provider coupling, but the evaluation suite is what makes switching safe.
Week eight validates the operating system
Deploy to staging with production-shaped integrations and permissions.
Run shadow traffic first. The agent processes real requests without affecting customers. Compare its route, answer, and tool path with the human outcome.
Then canary a small slice of eligible low-risk traffic. Keep a human approval or review path for uncertain and sensitive cases. Route users consistently to one behavior version during a conversation.
Verify:
- traces arrive with complete version metadata
- online evaluators and alerts run
- kill switches and rollback work
- cost budgets stop runaway execution
- support staff can inspect and report a bad result
- incident ownership reaches the right team
Week eight opens the production evidence loop. Evaluation continues for as long as the system runs.
Track the economics of resolved work
Cost per model call is an implementation metric. Cost per resolved eligible case is the operating metric.
Include model calls, judge calls, embeddings, retrieval, external APIs, human review, and incident overhead. Compare that cost with the current process.
Track false deflection separately. A case is not resolved because the agent answered and the customer left. It may have pushed the user into another channel.
Useful business measures include repeat contact within seven days, human correction, time to resolution, and downstream errors. Choose measures that match the process instead of assuming conversation thumbs-up means business success.
Grow the dataset from production
Online signals should feed a reviewed queue of candidate examples.
negative feedback or evaluator alert
↓
trace triage
↓
confirmed failure and root cause
↓
dataset example with failure label
↓
targeted evaluator or deterministic check
↓
offline verification
The dataset needs an owner. Categorize examples by intent, risk, source, and failure class. Review duplicates and outdated policy cases. More rows do not automatically mean better coverage.
Use sampled suites on pull requests and the full suite on main or scheduled runs. Keep a fixed high-risk sentinel set in every gate.
The calendar is less important than the order of evidence
Some teams can complete this work in four weeks. Others need months because data access, policy, or integrations dominate the schedule.
The reusable sequence is:
business contract
↓
representative work and failure classes
↓
evaluation harness
↓
data, traces, and controls
↓
model comparison
↓
controlled rollout
Selecting the model in week seven is a memorable way to make the point. The model becomes easier to choose after the team can measure it in context.
That is also the difference between a PoC and a production program. The PoC proves one path can work. The program builds the evidence and controls needed for all the paths that will not.
Start the series with Evaluation is the specification, or read the broader overview in From AI PoC to production: the engineering system around the model.

