Production RAG is a data lifecycle, not a vector database
How to engineer ingestion, freshness, lineage, permissions, deletion, retrieval evaluation, and telemetry for dependable production RAG.

A bank updates its interest-rate policy. Customers receive the notification, open the support assistant, and ask what changed. The assistant answers from the previous policy.
The model did exactly what the system asked. Retrieval supplied an old document because the new version never reached the vector index.
This failure is easy to miss because the application remains healthy. Traces show the agent reading stale policy, while the missing step happened earlier in the ingestion pipeline. The new document was never embedded and indexed.
This is the part of RAG architecture that disappears in demos. The demo has a folder of documents and a script. Production has updates, deletions, permissions, partial failures, and downstream systems that need to agree on which version is current.
Separate answer data from tracking data
A useful data architecture separates answer data from tracking data.
Question data helps the system answer. It includes documents, records, tool results, metadata, embeddings, and model context.
Tracking data explains what happened. It includes traces, retrieval results, evaluator scores, feedback, latency, and cost.
The two meet at a versioned identifier.
{
"source_id": "interest-rate-policy",
"source_version": "2026-08-14",
"chunk_id": "interest-rate-policy:2026-08-14:17",
"index_version": "retail-policy-31",
"ingestion_run": "ing-8821"
}
The answer pipeline stores these fields with the chunk. The trace records them when retrieval occurs. An operator can then move from a wrong answer to the exact source and ingestion run.
Without that join, observability stops at "retriever returned chunk 17," which is not enough to diagnose why chunk 17 was stale.
Treat ingestion as a state machine
A production ingestion pipeline needs explicit states and recovery rules.
discovered
↓
fetched
↓
validated
↓
normalized
↓
chunked
↓
embedded
↓
indexed
↓
verified
↓
active
Do not mark a document current when it has only been fetched. The active version should change after indexing and verification complete.
Persist state per source version. If embedding succeeds for 98 of 100 chunks, the run is incomplete. Retrying only the two failed chunks is cheaper than rebuilding the document, but activation must wait until the version passes its completeness rule.
Use idempotent keys for every stage. A repeated update event should not create duplicate chunks or multiple active versions.
document key source ID + source version
chunk key document key + chunk ordinal + chunker version
embedding key chunk key + embedding model version
index key embedding key + index version
These keys also make targeted deletion possible.
Define freshness as an SLO
"The index updates regularly" is not an operating target.
Define freshness per source class. A product manual may tolerate a 24-hour delay. A fraud blocklist may need updates within minutes. Policy documents may require activation before the communication announcing them goes out.
Track several timestamps:
source_updated_at
source_observed_at
ingestion_started_at
index_activated_at
last_verified_at
The difference between source_updated_at and index_activated_at is propagation lag. Alert when it crosses the source SLO.
Also track sources that have not changed. A connector that has been silently broken for a week can report no failed jobs because it never discovered new versions. A heartbeat or reconciliation scan should verify that the source and index still agree.
Put access control before retrieval
An LLM should never be the component deciding whether a user may see a document.
Permission-aware retrieval starts with identity and constructs a constrained search:
authenticated principal
↓
groups, tenant, role, purpose
↓
allowed source IDs or security filters
↓
retrieval
↓
authorized chunks only
Filtering after retrieval is weaker. Sensitive content has already entered application memory, traces, caches, or model context by then.
Carry the authorization context into tool calls and record the policy decision in the trace. The retriever's service identity should not turn a user request into administrator access.
Be careful with derived content. A summary generated from three restricted documents inherits their access constraints. An embedding does not become public because humans cannot read it directly. Caches need the same tenant and permission boundaries as the underlying results.
Version every transformation
The same document can produce different retrieval behavior after a chunking or embedding change.
Store the versions of:
- parser and normalization rules
- chunker and overlap settings
- metadata extraction
- embedding model
- index configuration
- reranker
A source version plus an embedding model is still insufficient if the chunker changed. The system needs to reproduce which text reached the embedding call.
Run retrieval experiments before promoting a new index. Use a fixed query set with expected relevant documents and compare recall, rank, latency, and empty results. Then run answer-level groundedness on top. A better answer score can hide worse retrieval if the model compensates from prior knowledge.
Evaluate retrieval and generation separately
An end-to-end answer can fail in two places.
The retriever may not return enough evidence. Or the model may receive good evidence and still answer incorrectly. One answer score cannot tell you which component to change.
For retrieval, track:
recall at k
precision at k
mean reciprocal rank
empty-result rate
permission-filter drop rate
freshness violations
For generation, track answer correctness, groundedness, citation accuracy, completeness, and appropriate abstention.
Use span-level evaluators where possible. LangSmith evaluation concepts recommend breaking quality down across retrieval, model, tool, and formatting components instead of evaluating only the final response.
Hard negative queries matter. If two policies share similar language but apply to different products, include questions that distinguish them. Metadata filters and rerankers often matter more on these cases than another prompt edit.
Design deletion before launch
Teams build ingestion first and discover deletion when a customer, regulator, or legal team requests it.
For one source version, identify every derivative:
raw object
normalized document
chunks
embeddings
vector index records
keyword index records
caches
evaluation examples
traces containing content
backups
Some stores can delete immediately. Others expire through retention. The system needs a deletion status that records what completed, what remains, and why.
Tombstone deleted source IDs so a delayed ingestion event cannot recreate them. Run periodic reconciliation between the source registry and active index. Verify deletion with the same seriousness as ingestion.
Keep provenance with the answer
A citation URL is useful to a user. It is not enough for operations.
The system should retain a provenance record with source ID, version, chunk IDs, retrieval time, index version, and permission context. If the source later changes, the team can still explain which version supported the answer at the time.
Do not display a citation the user cannot open. If the answer is derived from an internal source, the UI should route through the same authorization layer used by retrieval.
When sources disagree, avoid asking the model to choose silently. Encode source precedence or route the conflict to a deterministic policy. At minimum, tell the model to surface the conflict and abstain from action.
Use retrieval as a monitored dependency
A vector store can be available while retrieval quality is broken.
Monitor infrastructure and semantics:
- connector success and source heartbeat
- propagation lag by source class
- document and chunk counts by version
- embedding and indexing failures
- query latency and empty results
- retrieval score distributions
- permission denials
- online groundedness and citation feedback
Sudden count changes are useful. If an index loses 30% of its chunks after a parser deployment, an availability check will stay green. A volume and version check can stop activation.
Canary a new index behind a router. Send a slice of traffic to both versions in shadow mode, compare retrieval, and keep the old index available for rollback. The application should reference an index alias or logical version rather than a hard-coded physical collection.
LangChain is the access layer, not the data strategy
LangChain retrievers and document abstractions make it easier to connect stores and compose retrieval with an agent. They do not decide source ownership, freshness SLOs, deletion policy, or tenant boundaries.
Use framework metadata to carry source and version fields through the chain. Trace retriever calls in LangSmith. Evaluate retrieved documents independently from the answer. Keep the ingestion state and authorization policy in systems that can enforce them without model cooperation.
That separation makes provider and framework changes less dangerous. The data contract remains stable while the application layer changes.
The stale answer starts before the prompt
When a production assistant answers from an old policy, prompt tuning is an attractive distraction. It may make the wording more cautious, but it cannot retrieve a document that never reached the index.
A dependable RAG system knows which source version is active, how long propagation took, who may retrieve it, and which answer used it. It can roll an index forward, roll it back, and remove a source with its derivatives.
The next article moves from data to coordination: Multi-agent systems: choose the coordination cost before the pattern

