Categories

Architecture

LLM caches: saving tokens without dropping quality

A caching proxy in front of a language model can cut the token bill significantly, but it introduces subtle risks if the design is not careful. Which cache types work in production, where the usual traps sit, and how to add them without degrading the experience.

Architecture

Inference routers: choosing a model based on the request

An inference router decides which model answers each incoming request, weighing cost, latency and how hard the request actually is. Well-built inference routers cut total token spend by 30 to 70 percent with no quality loss the user can perceive. Four patterns cover most cases: length, task type, an auxiliary classifier, and learned routing.

Architecture

TigerBeetle: a database built for financial transactions

TigerBeetle is a distributed database written in Zig, specialized in one specific kind of workload: high-volume double-entry accounting with strong consistency guarantees. It does not aim to replace Postgres; it aims to be the right tool when the problem is counting financial transactions at millions per second without subtle failures.

Architecture

Platform engineering: consolidation after the boom

After three years of expansion and an overheated ecosystem around the term, platform engineering enters 2025 in a consolidation phase. The internal platforms that survive are the ones that understood their real function; those that mistook the label for the solution are dismantling their teams or cutting them drastically.

Technology

Fly.io: deploying globally without complicating your life

Fly.io has spent years selling the idea that deploying an application across several regions should be almost as simple as pushing an image and writing one config line. After several real projects on the platform, here is an honest read on what it delivers, what is missing, and who it is worth choosing over more classic options.

Technology

Microsoft Garnet: a high-performance cache alternative

Garnet is the open-source cache server Microsoft Research published in March 2024. Written from scratch in .NET 8, it speaks the Redis wire protocol, so existing clients connect unchanged, and it stores data through a hybrid memory-and-disk backend called Tsavorite. Core-affinity threading is what lets the Garnet cache outrun Redis on many-core hardware.

Artificial Intelligence

Testing with AI: the determinism problem

AI testing breaks the assumption every automated suite was built on, because the same input no longer produces the same output. Anthropic documents that temperature 0.0 is still not fully deterministic, and OpenAI's seed parameter only promises mostly reproducible results. What works instead is a layered belt of checks that catches regressions without tripping over ordinary variance.

Architecture

Citus: scaling Postgres horizontally without leaving it

Citus is a Postgres extension that spreads tables across worker nodes while the cluster still looks like a single Postgres server to your application. The coordinator intercepts query planning and distributes the work. Picking the distribution key is the decision that matters most, since every later query inherits its consequences.

Architecture

SQLite in production: patterns that have aged well

SQLite in production is a sound choice for small and mid-sized web applications once WAL mode is enabled, since concurrency on typical web loads improves by roughly two orders of magnitude. Litestream streams the WAL to S3-compatible object storage for point-in-time restore, and NVMe disks on cheap VPS plans removed the old disk objection.