Categories

Jacar categories — explore the topics A rocket whose eyes follow your cursor.
Artificial Intelligence

Mature LLM-as-judge: when to trust and when not

Using an LLM to judge another LLM became widespread in 2024 and remains, in 2026, the only scalable way to evaluate qualitative quality in LLM systems. It is reliable when judge-human correlation exceeds 0.7 on 30 cases and gets recalibrated quarterly; below that threshold, do not trust the number.

Methodologies

AI-integrated DevOps tools in my daily flow

After fourteen months testing AI-integrated DevOps tools across several teams, the stack that stays is small: Claude Code, Cursor, and Aider for code; PagerDuty AIOps, Datadog Bits AI, and Grafana Assistant for alert triage; and OpenTofu with OPA for infrastructure generation bounded by policy rules.

Artificial Intelligence

AI agent incidents: recovery runbooks that work

AI agents fail in production, and what matters is how you respond in the first twenty minutes. This runbook covers severity classification, isolating before investigating, purging contaminated memory, communicating without inventing facts, and turning every incident into a regression test before closing it as done.

Artificial Intelligence

LLM red teaming: a practical playbook

LLM red teaming has gone from an esoteric activity to a mandatory practice. With the OWASP Agentic Top 10 and the CSA Agentic AI Red Teaming Guide converging on shared vocabulary, this is the operational playbook any team deploying agents needs to have.

Methodologies

RICE: a prioritization framework for product roadmaps

The RICE framework is a prioritization methodology created by Intercom that produces a score by combining four factors: Reach, Impact, Confidence, and Effort. It divides the product of the first three by the estimated effort in person-months, so it can compare unrelated initiatives using one objective number.

Artificial Intelligence

Prompt Engineering: From Trick to Mature Discipline

Prompt engineering has moved from viral tricks to a discipline with reproducible patterns: few-shot, chain-of-thought, and structured output with function calling. Teams treating prompts like code (versioned, tested, and monitored) get consistently better results than those who improvise.

Artificial Intelligence

Lessons from agents in production in 2025: summary for 2026

By late 2025, 57.3 percent of organizations had agents in production, up from 51 percent a year earlier, according to LangChain's survey of more than 1,300 professionals. Three failure modes dominate the postmortems: degenerative reasoning loops, hallucinated data in RAG systems, and silent misalignment between the request and the interpretation.

Architecture

Consolidated platform engineering: who wins and who gets stuck

Platform engineering worked where teams built on concrete, painful problems and offered golden paths developers actually wanted, run with a product mindset. It stalled where the output was an empty Backstage portal: technically correct, unvisited, solving no operational problem. Three years after the Gartner hype of 2023, that split separates the winners from the sunk cost.

Artificial Intelligence

FinOps for AI workloads in 2026: the real pain

FinOps for AI counts different units than classic cloud FinOps: tokens, calls, computed embeddings and GPU time, all of which scale nonlinearly with use. The costliest habit is sending everything to frontier models; 40 to 70 percent of those calls run on mid-tier models with no noticeable quality loss. Uncached RAG and self-recursing agents do the rest.

Artificial Intelligence

Agents that drive the computer: patterns that work

Sixteen months after Anthropic first shipped computer use, with browser-use, OpenAI Operator and Gemini Computer Use all pushing in parallel, agents that drive the browser and desktop have moved from demo to real workflows. Time to review which patterns survive when you run them daily in production.

Methodologies

Product discovery with AI: practices that stick

Two years in, AI helps product discovery in one place above all: synthesizing interview transcripts. Generating hypotheses without real data has failed repeatedly, and simulated users produce systematic false positives about adoption. The practices that stick keep a human doing the critical analysis, because AI amplifies a good process and speeds a bad one toward failure.

Methodologies

Carbon-aware scheduling by default: first balance

Carbon aware scheduling delivers savings in proportion to how much of your workload can move in time or geography. If under 20 percent of it is elastic, cluster-wide gains stay modest. Deferrable jobs do best, cutting carbon intensity 15 to 30 percent: nightly batch, model training, CI builds and tests, report generation, video rendering.

Methodologies

SRE with AI: dashboards that actually help

Among the AI features in SRE dashboards, alert correlation is the one with demonstrated value: it groups dependent alerts from a single incident, cutting time to acknowledge and fatigue during a crisis. Automatic incident summaries help too. The real problem was never a shortage of information, it was separating signal from noise and correlating scattered symptoms.

Artificial Intelligence

LLM guardrails: frameworks and their real cost

Guardrails frameworks promise to filter language-model inputs and outputs to block data leaks, harmful content, or hallucinations. After evaluating four of the most popular ones in production, I cover what they actually do, what latency and billing cost they add, and when they pay off over simpler controls.

Artificial Intelligence

AI agent observability: tools and what to instrument first

Agents that chain calls to models, tools and memory are hard to debug without instrumentation designed for them. After a long year running agents in production, I cover what to measure first, which standards are consolidating, and which costly mistakes are avoided by getting the traces right from the start.

Architecture

Platform engineering: consolidation after the boom

After three years of expansion and an overheated ecosystem around the term, platform engineering enters 2025 in a consolidation phase. The internal platforms that survive are the ones that understood their real function; those that mistook the label for the solution are dismantling their teams or cutting them drastically.

Artificial Intelligence

Testing with AI: the determinism problem

AI testing breaks the assumption every automated suite was built on, because the same input no longer produces the same output. Anthropic documents that temperature 0.0 is still not fully deterministic, and OpenAI's seed parameter only promises mostly reproducible results. What works instead is a layered belt of checks that catches regressions without tripping over ordinary variance.

Methodologies

Carbon-aware computing: now the default behavior

Four years ago it was an academic curiosity. Today, scheduling workloads by grid carbon intensity is a built-in option in Kubernetes, in several cloud provider services, and in CI tooling. We look at what genuinely changed and what is still more promise than practice.

Methodologies

User research in the age of generative AI

Generative AI helps user research most in transcription, where it reliably saves hours, and in early note synthesis and discussion guide drafting. It does not replace real participants: synthetic personas return plausible answers rather than the genuine surprises interviews produce. Verify every quote in a final deliverable against the original transcript before anyone acts on it.

Methodologies

Migrating SSH to post-quantum cryptography: a practical guide

OpenSSH added hybrid post-quantum key exchange with ML-KEM in version 9.9 and made it the default algorithm in 10.0. The question is no longer whether to migrate SSH to post-quantum, but how to do it without breaking old clients: enable the hybrid mode, keep a classical fallback, and verify with ssh -v that the active algorithm is the right one.

Artificial Intelligence

Computer Use in production: agents that drive the interface

Computer Use in production works today for narrow, repetitive interface tasks where a human still checks the result. Anthropic shipped it in October 2024 calling it experimental and error-prone, and nine months on that framing still holds: teams run it on real work by constraining scope, not by trusting it end to end.

Methodologies

Continuous profiling with eBPF in production

Continuous profiling with eBPF samples every process's execution stack every few milliseconds without touching the code, then stores the history so you can compare last week's performance with today's. The cost measured in production runs between 1% and 3% of CPU, and it pays off most in databases, API gateways and high-concurrency services.

Methodologies

The Site Reliability Workbook: patterns we still use

Seven years on, the Site Reliability Workbook still earns its place in small teams through a handful of patterns: SLOs set slightly below what you already achieve, so the error budget is real and negotiable; a 28 to 30 day rolling window; and blameless postmortems that drop punishment while keeping accountability.

Artificial Intelligence

Continuous evaluation of RAG: dashboards that actually matter

Continuous RAG evaluation catches the quiet kind of failure: the system never goes down, never returns errors, never trips a latency alert, it simply answers worse as the index, the model, and user questions drift. Track retrieval precision, effective recall, faithfulness to the retrieved context, answer relevance, and p99 latency.

Methodologies

VEX: filtering vulnerability noise with context

VEX, the Vulnerability Exploitability eXchange, is a structured way for a vendor to state whether a CVE listed in an SBOM actually affects a product. Log4Shell in December 2021 showed why it exists: countless Java applications carried that critical CVE while never loading the vulnerable class. VEX marks which vulnerabilities are not exploitable, so scanner noise becomes signal.

Methodologies

Semgrep: modern SAST in your pipeline

Semgrep has grown into one of the most pragmatic static analyzers in the ecosystem. A look at why it works where other SAST tools fail, and how to fit it into a pipeline without turning it into noise.

Artificial Intelligence

AI governance in enterprise: committees, policies, audits

AI governance in a company means a standing committee, written policies, a model and use-case inventory, risk assessment, and audits. The first provisions of the EU AI Act took effect on 2 February 2025, banning practices such as social scoring and requiring minimum AI literacy for staff. Fines reach 35 million euros or 7 percent of global turnover.

Methodologies

SLSA v1.0: a mature framework for the software supply chain

SLSA v1.0 splits software supply-chain security into three tracks (Build, Source, and Dependencies), of which only Build is stabilized, with three levels: L1, L2, and L3. If you build in GitHub Actions, reaching L2 with Sigstore-signed provenance takes a few hours and is the starting point I recommend to any team.

Artificial Intelligence

How to Evaluate a RAG System Without Fooling Yourself

Measuring RAG quality rigorously takes more than skimming a handful of answers: it requires objective metrics (faithfulness, relevance, context precision, and coverage), a golden set of hundreds of curated questions, and regular human validation of the LLM judge to avoid misleading conclusions.

Methodologies

Green Software Principles: A Checklist for Teams

Software is not immaterial: every request and database query consumes electricity with a carbon footprint. The Green Software Foundation encodes eight practical principles to reduce that footprint without rewriting systems. The result is a more efficient service, a lower cloud bill, and readiness for ESG regulation.

Artificial Intelligence

LLM Observability: Traces, Costs, and Quality

LLM applications need three distinct observability planes: prompt and response traces for debugging hallucinations, per-token and per-feature cost tracking, and response quality evaluation. Mature tools like Langfuse, LangSmith, and Helicone cover all three planes with specific instrumentation.