FinOps for AI counts different units than classic cloud FinOps: tokens, calls, computed embeddings and GPU time, all of which scale nonlinearly with use. The costliest habit is sending everything to frontier models; 40 to 70 percent of those calls run on mid-tier models with no noticeable quality loss. Uncached RAG and self-recursing agents do the rest.
Sixteen months after Anthropic first shipped computer use, with browser-use, OpenAI Operator and Gemini Computer Use all pushing in parallel, agents that drive the browser and desktop have moved from demo to real workflows. Time to review which patterns survive when you run them daily in production.
A selection of postmortems published between 2025 and 2026 by teams running AI systems in production reveals repeated patterns: guardrail failures, silent model drift, hidden vendor dependency, and a collection of near-misses worth distilling.
The AI startup correction is already measurable: down rounds turned from anecdote into a visible statistical pattern from Q4 2025, and selective layoffs cluster in sales, research and operations at companies that over-hired. Survivors share a concrete problem, a concrete segment, and AI costs the business model can absorb. Thin wrappers over commercial models suffer most.
Claude Haiku 4.5, released by Anthropic on October 15, 2025, performs close to Sonnet 4 on structured tasks for roughly a third of the price. It carries a 200K-token context window and tool use on par with Sonnet 4.5. Pairing it as a filter ahead of Sonnet cuts total cost several times over.
Knowledge graphs spent two decades waiting for their moment. With LLMs now bridging free text and formal ontology, and the GraphRAG pattern already mature, the technology is back in the spotlight. Time to look at why it finally fits and where it actually pays off.
After two years watching every product invent its own interface for talking to an agent, by January 2026 a stable design consensus is emerging about which patterns work, which do not, and what the average user already expects. Time to write down what has settled.
Six months after A2A landed at the Linux Foundation, and after several implementation cycles from Google, Microsoft, and open projects, what version 1 of the protocol means and whether it is safe to build on yet.
European sovereign AI discourse has spent three years fueling headlines, public investment, and interstate agreements. We are starting to see which part of the promise has real technical substance and what a technical team expecting alternatives outside the US ecosystem can actually count on.
With MCP solving the agent-to-tool layer, a parallel problem surfaces: how do two agents from different vendors communicate with each other. Google's Agent2Agent protocol, donated to the Linux Foundation in June 2025, tries to fill that gap with an open standard.
Phi-3 is Microsoft Research's family of small language models aimed at the edge, and it competes directly with Llama 3.2, Gemma 2 and Qwen 2.5. Phi-3-mini holds 3.8B parameters and, quantized to 4 bits, fits in about 2 GB, running on a CPU with a neural accelerator, an integrated GPU or an NPU.
Large language models have spent two years promising effortless documentation for code, APIs and architecture. After watching dozens of projects try it, clear patterns emerge for where it works and where it just becomes more debt.
Guardrails frameworks promise to filter language-model inputs and outputs to block data leaks, harmful content, or hallucinations. After evaluating four of the most popular ones in production, I cover what they actually do, what latency and billing cost they add, and when they pay off over simpler controls.
Agents that chain calls to models, tools and memory are hard to debug without instrumentation designed for them. After a long year running agents in production, I cover what to measure first, which standards are consolidating, and which costly mistakes are avoided by getting the traces right from the start.
A caching proxy in front of a language model can cut the token bill significantly, but it introduces subtle risks if the design is not careful. Which cache types work in production, where the usual traps sit, and how to add them without degrading the experience.
An inference router decides which model answers each incoming request, weighing cost, latency and how hard the request actually is. Well-built inference routers cut total token spend by 30 to 70 percent with no quality loss the user can perceive. Four patterns cover most cases: length, task type, an auxiliary classifier, and learned routing.
AI testing breaks the assumption every automated suite was built on, because the same input no longer produces the same output. Anthropic documents that temperature 0.0 is still not fully deterministic, and OpenAI's seed parameter only promises mostly reproducible results. What works instead is a layered belt of checks that catches regressions without tripping over ordinary variance.
The term Agent OS has spent a year gaining traction across research and product circles. It describes a layer that goes well beyond an agent library: request scheduling, context management, persistent memory, and isolation. A look at the real state of that concept.
Model Context Protocol turns ten months old since Anthropic's announcement, and it is no longer just a proposal: hundreds of servers, cross-vendor implementations and a public registry now back it. A look at what has worked, what is still weak, and why 2025 marks the shift from curiosity to basic infrastructure.
After months of rumors, OpenAI released GPT-5 in early August. The first weeks of real-world use show a picture less spectacular than the marketing suggested and more useful than many expected. It is worth separating what is genuinely new from what is merely incremental.
Since 2 August 2025 the EU AI Act obligations for general-purpose models, national authorities, and the penalty regime are enforceable. A practical look at what changes for those of us deploying AI in Europe.
Redis 8.2 ships vector search as a native data type. The real question is whether it replaces a dedicated engine like Qdrant, Weaviate, or pgvector on workloads with millions of vectors and tight latency budgets, or only works as a bonus on top of the cache you already run.
Small language models have become genuinely useful. Phi-3.5, Gemma 2, and Llama 3.2 fit on modest hardware and solve bounded tasks without reaching the cloud. A look at where they fit on the factory floor and when skipping the large model pays off.
RAG 2.0 means retrieval built from several sources at once rather than a single vector search: dense embeddings, lexical matching, and knowledge graphs that capture relationships between entities, with a reranking layer ordering the final candidates. The 2023 pattern of one vector database plus an LLM no longer describes what production systems actually do.
Computer Use in production works today for narrow, repetitive interface tasks where a human still checks the result. Anthropic shipped it in October 2024 calling it experimental and error-prone, and nine months on that framing still holds: teams run it on real work by constraining scope, not by trusting it end to end.
MCP clients are now built into the editor itself: VS Code, Zed, Cursor, and several Neovim forks ship native support. The agent picks up project context in place, with no copying of files into a separate chat window. The practical questions become which servers to keep enabled and how to scope their permissions.
AI agents are starting to earn a real place in continuous integration pipelines: reviewing diffs, proposing fixes, generating missing tests. Six months of real-world use to separate the patterns that work from the ones that end up costing more time than they save.
Gemini 2.5 Pro reached preview on 25 March 2025 and general availability at the end of June, alongside the cheaper, faster Gemini 2.5 Flash. Two things separate it from Gemini 2.0: a one-million-token context window that behaves stably to at least 500k, and multimodality that has left the demo stage behind.
Anthropic released Claude Opus 4 and Claude Sonnet 4 on 22 May 2025, the first major naming jump since the 3.5 series. Claude 4 reasons noticeably better over long programming tasks: multi-hour refactors that previously stalled without a human nudge now run further alone, and the family targets agentic, multi-step flows.
A year after chat stopped being the only acceptable way to talk to an agent, UI patterns built specifically for agent tasks are emerging. I go through the ones starting to stick and the ones that are just cycle fashion.
After more than a thousand community MCP servers appeared, the shortlist worth keeping stays small. The five official Anthropic servers in daily use here are filesystem, git, read-only Postgres, fetch, and memory. Servers wrapping Slack, email, or broad SaaS APIs open more attack surface than they repay, and tool-chaining between servers remains unsolved.
Prompt injection is the most common vulnerability in LLM applications, and many teams defend against it with filters that do not work. We review defense layers backed by evidence, what actually works, and what is security theater.
For a decade, knowledge graphs were an academic idea with few real use cases, held back by the cost of building and maintaining the schema. LLMs have changed that equation: they now extract entities automatically and help anchor answers, audit reasoning, and support agents without hallucinating.
In AI systems the real cost is not EC2 instances but input tokens in RAG and agents, chained tool calls, and frequent reindexing; those vectors, plus unattributed experimental spend, concentrate most of the monthly production bill.
Continuous RAG evaluation catches the quiet kind of failure: the system never goes down, never returns errors, never trips a latency alert, it simply answers worse as the index, the model, and user questions drift. Track retrieval precision, effective recall, faithfulness to the retrieved context, answer relevance, and p99 latency.
Crunchbase and CB Insights first-quarter data confirm that global startup funding has rebounded, but nearly all of the growth is concentrated in startups presenting themselves as AI. The rest of the ecosystem remains in correction.
The LLM wrappers that survived the 2022 to 2024 startup wave own something their underlying model cannot supply: proprietary data, network effects, or a workflow users already live inside. Everything else was a thin prompt layer over an API, and structurally bad margins killed it as inference costs rose with usage.
LLM agent security became a real incident category the moment assistants gained tool access. Agents that call tools expose a far wider attack surface than chatbots, and it now carries assigned CVEs, published audit reports, and its own OWASP Top 10. The spread of the Model Context Protocol and agent-driven corporate workflows built that surface in under 18 months.
AI agents have moved from a lab curiosity to serious SDKs from three major providers. A reflection on moving from the flashy demo to an internal use case that shifts a real, measurable metric.
Graph RAG layers an explicit graph of entities and relationships on top of retrieval, so a question can be answered by traversing connections rather than by vector similarity alone. Microsoft Research published the GraphRAG paper in April 2024 and open-sourced the code that July. Lighter variants such as LightRAG and HippoRAG followed, running on Neo4j or Memgraph.
The AI features Figma has rolled out since Config 2024 are changing how product design teams work. A look at what each feature delivers, what remains human work, and which habits are taking hold across teams.
AI governance in a company means a standing committee, written policies, a model and use-case inventory, risk assessment, and audits. The first provisions of the EU AI Act took effect on 2 February 2025, banning practices such as social scoring and requiring minimum AI literacy for staff. Fines reach 35 million euros or 7 percent of global turnover.
Claude 3.7 Sonnet, released by Anthropic on February 24, is a careful refinement rather than a generational jump. The same model answers in standard mode or in an extended thinking mode you switch on per request, trading tokens and latency for better results on hard problems. It also ships Claude Code, a command-line tool for programmers.
A year ago open weights were a gamble; today they are a real production option. I review what has worked, what has not, and how Llama, DeepSeek, Qwen, and Mistral are fitting into enterprise architectures that used to depend on closed APIs.
vLLM remains the reference engine for serving LLMs on GPU in 2025: automatic prefix caching sharply cuts latency for repeated prompts, speculative decoding speeds up large models, and multi-LoRA support lowers the cost of multi-tenant SaaS, though multi-GPU support and non-NVIDIA hardware remain weak points.
GraphRAG has been in real enterprise use for over a year: during indexing, an LLM builds a knowledge graph that answers global questions about a corpus well, precisely where classic RAG fails because no single chunk holds the full answer. Here I compare indexing costs, the cases where it pays off, and the hybrid pattern that teams have settled on.
Three years after RLHF became popular, the model-alignment field is far richer. A review of RLHF, DPO, and newer methods such as KTO and ORPO, with criteria for choosing between them.
Google released Gemma 2 in mid-2024, and it has since seen real production use. A look at how it competes in the open-model ecosystem, which sizes actually make sense, and where its adoption has settled in.
6 min2434.2
We use first- and third-party cookies to analyze site traffic. You can accept them, reject them, or configure your choice.
Learn more about cookies
Cookie preferences
NecessaryEssential for the site to work. Always on.
AnalyticsHelp us understand how the site is used (Google Analytics).