Canary Releases for Prompt and Model Changes
Shipping a prompt change to 100% of traffic and watching the dashboard is how a subtle regression becomes a full-scale incident. Canary it like any other risky change.
Shipping a prompt change to 100% of traffic and watching the dashboard is how a subtle regression becomes a full-scale incident. Canary it like any other risky change.
Demo RAG and enterprise RAG are different engineering problems. Access control, multi-tenancy, freshness, and cost show up the moment real employees start using it.
Rolling back a bad code deploy is well understood. Rolling back a bad prompt, model, or retrieval config change needs the same discipline, applied to different artifacts.
A standard service runbook assumes failures look like errors or downtime. Agent incidents often look like the system working perfectly and producing something wrong.
You can't promise 99.9% uptime for a system whose failure modes are still partly unknown. Here's how to set an SLA that's honest about that without refusing to commit to anything.
A metadata filter that's easy to forget in one code path is a data leak waiting to happen. Multi-tenant RAG isolation needs to be structural, not a convention.
Redacting PII at generation time is too late — it's already been embedded, indexed, and potentially logged. Redaction belongs at ingestion.
A vector index has no inherent concept of who's allowed to see a document. Retrieval-time access control has to be enforced explicitly, or it doesn't exist.
Shipping RAG without evaluation means you find out it's broken from angry users. Here's how to measure retrieval quality and answer faithfulness before that happens.
A working list of every distinct way a RAG pipeline has failed in production, organized by which stage caused it — closing out this stretch of RAG-architecture posts.