AI Engineering 133
- Benchmarking a Browser Agent Against a Human on the Same 30-Field Form
- The Manus Acquisition and What It Signals About Browser Agent Consolidation
- Inside an Agentic Browser: How Form-Filling and Navigation Actually Work
- Computer-Use Agents in 2026: What Actually Changed Since the First Demos
- Building a Business Case for a Vertical Agent, Not a Platform
- IT Operations Agents: Ticket Triage as the Proving Ground for Enterprise Trust
- Procurement and Document-Heavy Workflows: The Agent Use Case Nobody's Excited About and Everyone's Shipping
- Multi-Vendor Model Routing Inside a Single Orchestrated Workflow
- The Orchestration Layer: Where 2026's Enterprise Agent Value Actually Concentrates
- From Chat to Execution: Measuring Agents by What They Close, Not What They Say
- Why Narrow, Vertical Agents Are Winning Over General-Purpose Assistants in 2026
- The LLMOps Maturity Model: Where Your Team Actually Stands
- Closing the Loop: From Week One's Architecture to a Mature Platform
- A Roadmap Template for an Agent Platform's Next Two Quarters
- Model Routing and Cascades: Cutting LLM Costs Without Losing Quality
- Six Months In: What We'd Do Differently
- Migrating from Prototype to Platform Without a Rewrite
- Buy vs Build for an Internal Agent Platform
- Capacity Planning for Unpredictable Agent Workloads
- Chargeback Models for Shared AI Infrastructure Spend
- Standing Up an Agent Governance Council
- TokenOps: A FinOps Practice for LLM and Agent Cost Management
- Org Design for a Platform Team Supporting Agents
- Setting SLOs for an Agentic System That Has Never Had One
- Multi-Provider Redundancy Without Doubling Your Bill
- Fallback Chains for Provider Outages
- Routing Policies: A Deeper Look at the Decision Logic
- Evaluating Cheaper Models Without Quietly Losing Quality
- Defending Against Memory and Context Poisoning in Long-Running Agents
- Small Model Distillation as a Cost Lever
- Batching Requests for Cost Savings Without Hurting Latency
- Caching Strategies That Meaningfully Cut Token Spend
- Catching Cost Anomalies in LLM Spend Before the Invoice
- Building a Token Budget Dashboard Engineers Actually Check
- An Incident Response Runbook for Agent Security Breaches
- The OWASP Agentic AI Top 10: A Field Guide to Threats Beyond Prompt Injection
- Insider-Threat Scenarios Unique to Agentic Systems
- Supply-Chain Risk in the Agent Tool Ecosystem
- Defense in Depth Against Prompt Injection
- Red-Teaming Your Own Agents, On a Schedule
- Threat Modeling an Agentic System Before It Ships
- A Reference Architecture for Agent Infrastructure, Assembled
- Structured Outputs and Tool-Call Contracts: The Reliability Layer Agents Actually Need
- Multi-Region Deployment for Latency-Sensitive Agents
- Secrets Management for Agents That Call Real APIs
- Rate-Limiting Shared Infrastructure Fairly Across Agents
- Building an Agent Gateway in Front of Shared Tool Servers
- Negotiation Protocols Between Autonomous Agents
- Observability for MCP Calls: What to Trace
- A2A and the Multi-Agent Mesh: Interoperability Beyond a Single Framework
- The Infra Cost of Long Context, Measured
- Validating JSON Schema at the Edge, Before It Reaches the Agent
- Streaming Structured Outputs Without Corrupting the Schema
- Retries and Idempotency for Tool Calls
- Recovering from Function-Calling Errors Gracefully
- Evolving a Structured Output Schema Without Breaking Consumers
- Building Production-Grade Agent Memory: Tiers, Write Policies, and Retrieval
- Cross-Org Agent Handoffs: Where Interoperability Gets Hard
- Trust and Identity Between Agents That Don't Share an Owner
- Designing an A2A Agent Card That Tells the Truth
- Vector Memory vs Graph Memory: Different Failure Modes
- Resolving Write Conflicts in Shared Agent Memory
- Episodic vs Semantic Memory in Practice
- MCP in Production: Wiring Agents to Tools and Data at Enterprise Scale
- Memory Eviction Policies: What an Agent Should Forget
- Building an Internal MCP Server Registry
- Versioning MCP Tools Without Breaking Existing Agents
- Tool Discovery at Scale Across Dozens of MCP Servers
- Auth Patterns for MCP Servers in Production
- Prompt Caching Strategies That Actually Move the Cost Needle
- Context Engineering: The Discipline That's Replacing Prompt Engineering
- Budgeting a Context Window Like It's a Scarce Resource
- Three Months of Spec-Driven Development: A Retrospective
- Generating Documentation from Specs Instead of Code Comments
- Spec-Driven Database Migrations
- Spec-Driven API Design Before Writing a Single Endpoint
- Spec-Driven Infrastructure-as-Code
- From Vibe Coding to Spec-Driven Development: A Migration Playbook
- Spec Failures: Three Case Studies and What Went Wrong
- Pair Programming with an Agent Under a Spec-Driven Workflow
- Spec-Driven Development vs Test-Driven Development
- GitHub Spec-Kit vs Kiro: Two Approaches to Spec-Driven Development Compared
- Walking Through Kiro's Requirements Notation
- Real Spec-Kit Constitution Examples, Annotated
- Kiro IDE: Specs, Steering, and Hooks for Production-Grade Agentic Coding
- Keeping Specs Consistent Across a Multi-Repo Codebase
- Spec-Driven Refactors of Code Nobody Wants to Touch
- Handling Genuine Ambiguity in a Spec
- Code Review for Spec-Driven Changes: What to Look For
- Wiring Lint and Test Hooks into an Agentic IDE Workflow
- Kiro Steering Files: A Deep Dive
- GitHub Spec-Kit: A Practical Guide to Spec-Driven Development with Coding Agents
- A Starter Library of Spec Templates
- Measuring SDD Adoption Across a Team
- Onboarding New Engineers with Spec-Driven Workflows
- Version-Controlling Specs Alongside Code
- Agent-Authored Specs vs Human-Authored: Who Should Write the First Draft?
- Writing Specs for a Legacy Migration, Not Just Greenfield Code
- Vibe Coding vs. Spec-Driven Development: Why Agentic Coding Agents Need Specs
- Spec-Driven Testing: Deriving Test Cases from the Spec Itself
- Detecting Spec Drift Before It Becomes Tech Debt
- A Spec Review Checklist for Agentic Coding Sessions
- Writing a Constitution Doc an Agent Will Actually Follow
- Guardrails for LLM Agents: Input, Output, and Action Validation
- Building an LLM-as-Judge You Can Actually Trust
- LangGraph in Production: Patterns for Reliable Multi-Step Agents
- The Performance Overhead of Guardrails, Measured
- Testing Guardrails Like You'd Test Any Other Unit
- Tuning Guardrails to Cut False Positives Without Opening Holes
- Wiring Up Ragas: A Hands-On Guide to RAG Evaluation
- Kill Switches: Designing the Agent's Emergency Stop
- Layered Defense: Why One Guardrail Is Never Enough
- Evaluating Agent Skills: A Framework for Measuring What Matters
- Guardrails for Streaming Responses
- Jailbreak Defense Patterns That Hold Up Under Testing
- PII Guardrails: Catching Leaks Before They Leave the System
- Output Guardrails for Structured Data, Not Just Text
- Red-Teaming Your Own Guardrails Before Someone Else Does
- Designing a LangGraph State Schema That Scales
- Agent Skill Design Patterns: What Works in Production
- Human-in-the-Loop Interrupts in LangGraph
- LangGraph Checkpointing: A Deep Dive
- The Hidden Cost of Running Your Own Evaluation Suite
- Human-in-the-Loop Evaluation: When Automated Scoring Isn't Enough
- Gating Deploys on Eval Regressions in CI
- Curating a Golden Dataset for Agent Evaluation
- Building Your First Agent Skill: From Definition to Production
- Deprecating a Skill Without Breaking Downstream Agents
- Skill Composition Patterns: Small Skills vs One Big One
- Building a Test Harness for Agent Skills
- Skill Discovery at Scale: When an Agent Has 200 Skills to Choose From
- Skill Versioning: Shipping Changes Without Breaking Callers
- What Are Agent Skills? A Practical Introduction