Kill Switches: Designing the Agent's Emergency Stop
When every layered defense misses something in production, the last resort is a kill switch that actually works under pressure — not one that's untested until the day it's needed.
When every layered defense misses something in production, the last resort is a kill switch that actually works under pressure — not one that's untested until the day it's needed.
A synthesis of the last several posts' recurring theme — every effective guardrail pattern covered so far has been a layered one. Here's why that keeps being true.
I promised a follow-up on skill evaluation in the last post. Here's the three-layer framework I actually use to know whether a skill library is working.
The same tension covered earlier for RAG grounding — streaming vs. pre-validation — applies to guardrails generally. Here's how to apply it across the full guardrail suite, not just groundedness.
A single-layer jailbreak defense — one classifier, one keyword list — is exactly the kind of thing red-teaming reliably finds a way around. Layered defense holds up where single defenses don't.
Redacting PII at ingestion (covered earlier in this blog) stops indexed PII from being retrievable. It doesn't stop a model from generating or repeating PII that reached it through some other path — that needs an output-side guardrail too.
Most guardrail tooling is built to scan prose. A structured JSON or function-call output needs guardrails that understand the schema, not just the text.
A guardrail that's never been attacked hasn't been tested — it's been assumed. Red-teaming your own system finds the gaps before a real adversary or a curious user does.
A state schema that starts as one flat dictionary works fine for a three-node graph and becomes unmanageable well before twenty. Here's how to structure it so it doesn't.
After building dozens of agent skills across different domains, some patterns consistently work and some consistently don't. Here's what I've learned.