Guardrails frameworks promise to filter language-model inputs and outputs to block data leaks, harmful content, or hallucinations. After evaluating four of the most popular ones in production, I cover what they actually do, what latency and billing cost they add, and when they pay off over simpler controls.
Agents that chain calls to models, tools and memory are hard to debug without instrumentation designed for them. After a long year running agents in production, I cover what to measure first, which standards are consolidating, and which costly mistakes are avoided by getting the traces right from the start.
A caching proxy in front of a language model can cut the token bill significantly, but it introduces subtle risks if the design is not careful. Which cache types work in production, where the usual traps sit, and how to add them without degrading the experience.
An inference router decides which model answers each incoming request, weighing cost, latency and how hard the request actually is. Well-built inference routers cut total token spend by 30 to 70 percent with no quality loss the user can perceive. Four patterns cover most cases: length, task type, an auxiliary classifier, and learned routing.
AI testing breaks the assumption every automated suite was built on, because the same input no longer produces the same output. Anthropic documents that temperature 0.0 is still not fully deterministic, and OpenAI's seed parameter only promises mostly reproducible results. What works instead is a layered belt of checks that catches regressions without tripping over ordinary variance.
The term Agent OS has spent a year gaining traction across research and product circles. It describes a layer that goes well beyond an agent library: request scheduling, context management, persistent memory, and isolation. A look at the real state of that concept.
Model Context Protocol turns ten months old since Anthropic's announcement, and it is no longer just a proposal: hundreds of servers, cross-vendor implementations and a public registry now back it. A look at what has worked, what is still weak, and why 2025 marks the shift from curiosity to basic infrastructure.
After months of rumors, OpenAI released GPT-5 in early August. The first weeks of real-world use show a picture less spectacular than the marketing suggested and more useful than many expected. It is worth separating what is genuinely new from what is merely incremental.
Since 2 August 2025 the EU AI Act obligations for general-purpose models, national authorities, and the penalty regime are enforceable. A practical look at what changes for those of us deploying AI in Europe.
Redis 8.2 ships vector search as a native data type. The real question is whether it replaces a dedicated engine like Qdrant, Weaviate, or pgvector on workloads with millions of vectors and tight latency budgets, or only works as a bonus on top of the cache you already run.
Small language models have become genuinely useful. Phi-3.5, Gemma 2, and Llama 3.2 fit on modest hardware and solve bounded tasks without reaching the cloud. A look at where they fit on the factory floor and when skipping the large model pays off.
RAG 2.0 means retrieval built from several sources at once rather than a single vector search: dense embeddings, lexical matching, and knowledge graphs that capture relationships between entities, with a reranking layer ordering the final candidates. The 2023 pattern of one vector database plus an LLM no longer describes what production systems actually do.
Computer Use in production works today for narrow, repetitive interface tasks where a human still checks the result. Anthropic shipped it in October 2024 calling it experimental and error-prone, and nine months on that framing still holds: teams run it on real work by constraining scope, not by trusting it end to end.
MCP clients are now built into the editor itself: VS Code, Zed, Cursor, and several Neovim forks ship native support. The agent picks up project context in place, with no copying of files into a separate chat window. The practical questions become which servers to keep enabled and how to scope their permissions.
AI agents are starting to earn a real place in continuous integration pipelines: reviewing diffs, proposing fixes, generating missing tests. Six months of real-world use to separate the patterns that work from the ones that end up costing more time than they save.
Gemini 2.5 Pro reached preview on 25 March 2025 and general availability at the end of June, alongside the cheaper, faster Gemini 2.5 Flash. Two things separate it from Gemini 2.0: a one-million-token context window that behaves stably to at least 500k, and multimodality that has left the demo stage behind.
Anthropic released Claude Opus 4 and Claude Sonnet 4 on 22 May 2025, the first major naming jump since the 3.5 series. Claude 4 reasons noticeably better over long programming tasks: multi-hour refactors that previously stalled without a human nudge now run further alone, and the family targets agentic, multi-step flows.
A year after chat stopped being the only acceptable way to talk to an agent, UI patterns built specifically for agent tasks are emerging. I go through the ones starting to stick and the ones that are just cycle fashion.
After more than a thousand community MCP servers appeared, the shortlist worth keeping stays small. The five official Anthropic servers in daily use here are filesystem, git, read-only Postgres, fetch, and memory. Servers wrapping Slack, email, or broad SaaS APIs open more attack surface than they repay, and tool-chaining between servers remains unsolved.
Prompt injection is the most common vulnerability in LLM applications, and many teams defend against it with filters that do not work. We review defense layers backed by evidence, what actually works, and what is security theater.
For a decade, knowledge graphs were an academic idea with few real use cases, held back by the cost of building and maintaining the schema. LLMs have changed that equation: they now extract entities automatically and help anchor answers, audit reasoning, and support agents without hallucinating.
In AI systems the real cost is not EC2 instances but input tokens in RAG and agents, chained tool calls, and frequent reindexing; those vectors, plus unattributed experimental spend, concentrate most of the monthly production bill.
Continuous RAG evaluation catches the quiet kind of failure: the system never goes down, never returns errors, never trips a latency alert, it simply answers worse as the index, the model, and user questions drift. Track retrieval precision, effective recall, faithfulness to the retrieved context, answer relevance, and p99 latency.
Crunchbase and CB Insights first-quarter data confirm that global startup funding has rebounded, but nearly all of the growth is concentrated in startups presenting themselves as AI. The rest of the ecosystem remains in correction.
The LLM wrappers that survived the 2022 to 2024 startup wave own something their underlying model cannot supply: proprietary data, network effects, or a workflow users already live inside. Everything else was a thin prompt layer over an API, and structurally bad margins killed it as inference costs rose with usage.
LLM agent security became a real incident category the moment assistants gained tool access. Agents that call tools expose a far wider attack surface than chatbots, and it now carries assigned CVEs, published audit reports, and its own OWASP Top 10. The spread of the Model Context Protocol and agent-driven corporate workflows built that surface in under 18 months.
AI agents have moved from a lab curiosity to serious SDKs from three major providers. A reflection on moving from the flashy demo to an internal use case that shifts a real, measurable metric.
Graph RAG layers an explicit graph of entities and relationships on top of retrieval, so a question can be answered by traversing connections rather than by vector similarity alone. Microsoft Research published the GraphRAG paper in April 2024 and open-sourced the code that July. Lighter variants such as LightRAG and HippoRAG followed, running on Neo4j or Memgraph.
The AI features Figma has rolled out since Config 2024 are changing how product design teams work. A look at what each feature delivers, what remains human work, and which habits are taking hold across teams.
AI governance in a company means a standing committee, written policies, a model and use-case inventory, risk assessment, and audits. The first provisions of the EU AI Act took effect on 2 February 2025, banning practices such as social scoring and requiring minimum AI literacy for staff. Fines reach 35 million euros or 7 percent of global turnover.
Claude 3.7 Sonnet, released by Anthropic on February 24, is a careful refinement rather than a generational jump. The same model answers in standard mode or in an extended thinking mode you switch on per request, trading tokens and latency for better results on hard problems. It also ships Claude Code, a command-line tool for programmers.
A year ago open weights were a gamble; today they are a real production option. I review what has worked, what has not, and how Llama, DeepSeek, Qwen, and Mistral are fitting into enterprise architectures that used to depend on closed APIs.
vLLM remains the reference engine for serving LLMs on GPU in 2025: automatic prefix caching sharply cuts latency for repeated prompts, speculative decoding speeds up large models, and multi-LoRA support lowers the cost of multi-tenant SaaS, though multi-GPU support and non-NVIDIA hardware remain weak points.
GraphRAG has been in real enterprise use for over a year: during indexing, an LLM builds a knowledge graph that answers global questions about a corpus well, precisely where classic RAG fails because no single chunk holds the full answer. Here I compare indexing costs, the cases where it pays off, and the hybrid pattern that teams have settled on.
Three years after RLHF became popular, the model-alignment field is far richer. A review of RLHF, DPO, and newer methods such as KTO and ORPO, with criteria for choosing between them.
Google released Gemma 2 in mid-2024, and it has since seen real production use. A look at how it competes in the open-model ecosystem, which sizes actually make sense, and where its adoption has settled in.
o3-mini, the first public release of OpenAI's o3 reasoning series, clearly improves logic, math, and complex code over GPT-4o, though it answers slower and still hallucinates facts. This analysis, based on weeks of real use, explains where it pays off and where it does not.
Gemini 2.0, announced by Google in December, puts tool use and agent behavior at the center of the product: it is designed to execute actions, not only generate text. Its clearest advantages are Flash's one-million-token context window, cheap input tokens, and first-class access to Search, Maps, Cloud, and Workspace. Claude 3.5 Sonnet still leads on complex reasoning.
Two years running AI-assisted code review in a real team leave a clear balance: AI catches mechanical oversights well and writes useful pull-request summaries, but it struggles with architectural judgment and produces many false positives on subtle bugs. The single decision that helped the most was not blocking merges on its automated comments.
Meta publicó Llama 3.2 con modelos tan pequeños como 1B y 3B, pensados específicamente para ejecutarse en dispositivos. Análisis de qué pueden hacer realmente y cómo se comparan con las alternativas.
Qualcomm, Intel and AMD Copilot+ processors have normalised the presence of an NPU in everyday PCs. A 40 TOPS NPU can run quantised Phi-3 Mini drawing just 5-10 W, versus 40-50 W for a laptop GPU doing the same task. What actually changes for running AI models locally, and when it is worth it.
Measuring RAG quality rigorously takes more than skimming a handful of answers: it requires objective metrics (faithfulness, relevance, context precision, and coverage), a golden set of hundreds of curated questions, and regular human validation of the LLM judge to avoid misleading conclusions.
Hybrid search combines BM25 and vector retrieval to cover what each misses alone. Vectors fail on exact identifiers like SKUs or CVEs; BM25 fails when query and document use different vocabulary for the same idea. Reciprocal Rank Fusion (RRF) merges both rankings without depending on their score scales.
llama.cpp is the C++ library that powers Ollama and much of the local-LLM ecosystem. 2024 added speculative decoding with two- to three-fold speedups, an RPC server for sharding layers across machines, and a stable GGUF format. Ollama covers 90% of cases; going direct pays off with uncommon hardware or specific flags.
Ollama became the standard for running large language models locally in 2024. It wraps llama.cpp in a single binary with Docker-style CLI and an OpenAI-compatible API. Phi-3 Mini runs in 4 GB; Llama 3.1 8B Q4 needs 6 GB. For production traffic at scale, vLLM remains the correct choice.
Model Context Protocol (MCP) is the open standard Anthropic published on 25 November 2024 to connect language models with external data and tools over JSON-RPC 2.0. It does not replace function calling: it standardises the server side, aiming to become for context what the Language Server Protocol is for code editors.
Product-market fit for LLM-powered products still depends on the same classic signals: cohort retention, NPS, and revenue expansion. What changes are the higher quality baseline, faster competitor iteration, and where durable moats come from: proprietary data, workflow integration, and network effects.
4 min2934.2
We use first- and third-party cookies to analyze site traffic. You can accept them, reject them, or configure your choice.
Learn more about cookies
Cookie preferences
NecessaryEssential for the site to work. Always on.
AnalyticsHelp us understand how the site is used (Google Analytics).