Docker Agent is a Docker CLI plugin for building and running AI agents from a declarative YAML file, no code required: you define one or more agents, give them MCP tools, run them with docker agent run, and share them through any OCI registry, on OpenAI, Anthropic, Gemini or other providers.
Open GSD (Git. Ship. Done.) is an open-source, MIT-licensed toolkit for steering coding agents without losing context: it splits work into five phases (discuss, plan, execute, verify and ship) and delegates the heavy lifting to subagents that each start with a clean context. Its core is the gsd-core engine and the gsd-pi terminal agent.
Skills package reusable capabilities; subagents isolate bounded-task execution. Together they form the most effective pattern for composing complex agents in 2026.
Using an LLM to judge another LLM became widespread in 2024 and remains, in 2026, the only scalable way to evaluate qualitative quality in LLM systems. It is reliable when judge-human correlation exceeds 0.7 on 30 cases and gets recalibrated quarterly; below that threshold, do not trust the number.
Opus 4.7 launched as Anthropic's most capable model, with emphasis on long-horizon agentic work. After two months of intensive use, these are the practical changes versus Opus 4.6.
The first invoice for a production agent usually runs double or triple the estimate. This article walks through five real levers, in priority order, caching, routing, context control, batching, and telemetry, to cut cost without touching perceived quality.
AI agents fail in production, and what matters is how you respond in the first twenty minutes. This runbook covers severity classification, isolating before investigating, purging contaminated memory, communicating without inventing facts, and turning every incident into a regression test before closing it as done.
LLM red teaming has gone from an esoteric activity to a mandatory practice. With the OWASP Agentic Top 10 and the CSA Agentic AI Red Teaming Guide converging on shared vocabulary, this is the operational playbook any team deploying agents needs to have.
Reliable agents come from measurement, not from better models or prompts. A production agent evaluation setup starts with a golden dataset of 30 to 200 cases, split roughly 60 percent normal usage, 30 percent edge cases and 10 percent adversarial, and it never uses the same model as both worker and judge.
An Agent OS is a runtime layer built to run AI agents rather than ordinary applications, and after six months of production deployments the trade-off is clear. A dedicated agent stack starts slower but stays stable; Kubernetes with orchestration bolted on top moves faster early, then hits observability and policy limits. It pays off from five active agents.
After two years of pilots and a year of agents in production, governance has moved from an aspirational committee to an operational control. What audits ask for, what broke in 2025, and which guardrails absorb most incidents.
By late 2025, 57.3 percent of organizations had agents in production, up from 51 percent a year earlier, according to LangChain's survey of more than 1,300 professionals. Three failure modes dominate the postmortems: degenerative reasoning loops, hallucinated data in RAG systems, and silent misalignment between the request and the interpretation.
Twenty months after the initial announcement, Model Context Protocol went from curiosity to de-facto standard among agent clients and servers. What is available, which servers are worth it, which problems remain open, and how it compares to earlier protocol maps.
Claude Haiku 4.5, released by Anthropic on October 15, 2025, performs close to Sonnet 4 on structured tasks for roughly a third of the price. It carries a 200K-token context window and tool use on par with Sonnet 4.5. Pairing it as a filter ahead of Sonnet cuts total cost several times over.
After two years watching every product invent its own interface for talking to an agent, by January 2026 a stable design consensus is emerging about which patterns work, which do not, and what the average user already expects. Time to write down what has settled.
Six months after A2A landed at the Linux Foundation, and after several implementation cycles from Google, Microsoft, and open projects, what version 1 of the protocol means and whether it is safe to build on yet.
With MCP solving the agent-to-tool layer, a parallel problem surfaces: how do two agents from different vendors communicate with each other. Google's Agent2Agent protocol, donated to the Linux Foundation in June 2025, tries to fill that gap with an open standard.
Agents that chain calls to models, tools and memory are hard to debug without instrumentation designed for them. After a long year running agents in production, I cover what to measure first, which standards are consolidating, and which costly mistakes are avoided by getting the traces right from the start.
Computer Use in production works today for narrow, repetitive interface tasks where a human still checks the result. Anthropic shipped it in October 2024 calling it experimental and error-prone, and nine months on that framing still holds: teams run it on real work by constraining scope, not by trusting it end to end.
MCP clients are now built into the editor itself: VS Code, Zed, Cursor, and several Neovim forks ship native support. The agent picks up project context in place, with no copying of files into a separate chat window. The practical questions become which servers to keep enabled and how to scope their permissions.
AI agents are starting to earn a real place in continuous integration pipelines: reviewing diffs, proposing fixes, generating missing tests. Six months of real-world use to separate the patterns that work from the ones that end up costing more time than they save.
A year after chat stopped being the only acceptable way to talk to an agent, UI patterns built specifically for agent tasks are emerging. I go through the ones starting to stick and the ones that are just cycle fashion.
After more than a thousand community MCP servers appeared, the shortlist worth keeping stays small. The five official Anthropic servers in daily use here are filesystem, git, read-only Postgres, fetch, and memory. Servers wrapping Slack, email, or broad SaaS APIs open more attack surface than they repay, and tool-chaining between servers remains unsolved.
Y Combinator's W25 and S25 cohorts show a historic tilt toward vertical agents and developer tools, with outcome-based pricing emerging as a new model. I break down the visible patterns, the business models on display, and what founders operating outside Silicon Valley should copy from this reading of the batch.
6 min7974.3
We use first- and third-party cookies to analyze site traffic. You can accept them, reject them, or configure your choice.
Learn more about cookies
Cookie preferences
NecessaryEssential for the site to work. Always on.
AnalyticsHelp us understand how the site is used (Google Analytics).