Skill Discovery at Scale: When an Agent Has 200 Skills to Choose From
Stuffing 200 tool schemas into a single prompt degrades selection accuracy and burns context. Skill discovery needs its own retrieval step past a certain scale.
Stuffing 200 tool schemas into a single prompt degrades selection accuracy and burns context. Skill discovery needs its own retrieval step past a certain scale.
An agent skill's interface is a contract with every agent that calls it. Changing it without versioning breaks callers you may not even know exist.
Agent skills are the modular building blocks that give LLMs the ability to act. This post breaks down what they are, how they work, and why getting the design right matters.
"Production-ready" gets used as a vague vibe check. This closes out the reliability stretch of this series with a concrete, checkable definition.
A standard postmortem template asks questions built for deterministic software failures. Agent incidents need a few additions, or the document misses exactly the parts worth learning from.
Standard load testing tools assume deterministic, fast responses. An agent pipeline is neither, and that breaks more of the standard load-testing playbook than expected.
A rules-based chatbot and an agentic replacement fail differently enough that a straight cutover is risky. The safer path runs both and compares before retiring the old one.
A checklist of practices that separate agentic systems that survive contact with real users from the ones that get quietly turned off after a bad week.
An agent that never escalates isn't more autonomous — it's more likely to confidently do the wrong thing. Escalation design is a first-class architecture decision, not a fallback bolted on.
A code-executing agent tool without a sandbox is a code execution vulnerability with extra steps. Here's what an actual production-safe sandbox needs.