Human-in-the-Loop Interrupts in LangGraph
LangGraph's interrupt mechanism, combined with checkpointing, is what makes a genuinely durable approval pause possible — one that survives a process restart while it's waiting on a human.
LangGraph's interrupt mechanism, combined with checkpointing, is what makes a genuinely durable approval pause possible — one that survives a process restart while it's waiting on a human.
LangGraph's checkpointer is what makes durable, resumable agent state possible — and picking the wrong backend for your use case is an easy, costly mistake.
An LLM-as-judge eval suite is itself a set of LLM calls, and at real CI frequency, that cost adds up in ways teams don't budget for until the invoice arrives.
Automated eval scales; it also can't catch everything. Knowing which gap human review actually fills — and how to run it without it becoming a bottleneck — matters more than picking one approach over the other.
Running an eval suite manually before a big change catches obvious regressions. Wiring it into CI as a hard gate catches the small ones nobody thought to check for.
A golden dataset built entirely from imagined test cases misses the failure modes real usage actually produces. Here's a practical process for building one that reflects reality.
A hands-on walkthrough of building a real agent skill in Python — covering interface design, execution, error handling, and output formatting for the model.
Removing a skill is riskier than removing a traditional API endpoint, because the callers referencing it aren't just code — they're prompts, few-shot examples, and model training data you may not fully control.
The atomic-vs-composite tradeoff from earlier in this series has recognizable patterns once you've built enough skills to see them repeat. A deeper look at when each pattern actually wins.
A skill needs two distinct kinds of tests — does the function work, and does the model actually invoke it correctly — and most teams only build the first kind.