Detecting Spec Drift Before It Becomes Tech Debt
A spec that described the system accurately at merge time slowly stops matching reality as follow-up changes land without updating it. That gap is spec drift, and it compounds silently.
A spec that described the system accurately at merge time slowly stops matching reality as follow-up changes land without updating it. That gap is spec drift, and it compounds silently.
Reviewing a spec before an agent implements it is far cheaper than reviewing the resulting code — but only if the review checks for the failure modes specs actually have.
A constitution doc that reads like a mission statement gets ignored by a coding agent the same way vague instructions always do. Here's what makes one actually enforceable.
Guardrails deserve more than a bullet point. A layered look at input validation, output validation, and action-level gates — and why none of them belong solely in the system prompt.
LLM judges fail in predictable, well-documented ways. Here's how I design rubrics, calibrate against human labels, and know when a judge score is actually trustworthy.
The patterns that separate a LangGraph prototype from a graph that survives real production traffic — checkpointing, interrupts, subgraphs, and streaming.
Every guardrail layer covered this month adds latency. Closing out the guardrails stretch of this blog with the actual numbers, not just the architectural case for having them.
Guardrails often get validated ad hoc — try a few known bad inputs, confirm they're blocked. A real test suite, with the same rigor as application code, catches far more.
A guardrail tuned purely to minimize false negatives frustrates legitimate users constantly. One tuned purely to minimize false positives misses real issues. Here's how to actually navigate that tradeoff.
The theory of RAG evaluation is one thing. Actually wiring Ragas into your own pipeline — judge model config, dataset construction, debugging a low score — is where most of the friction lives.