Categories

Artificial Intelligence

Mature LLM-as-judge: when to trust and when not

Using an LLM to judge another LLM became widespread in 2024 and remains, in 2026, the only scalable way to evaluate qualitative quality in LLM systems. It is reliable when judge-human correlation exceeds 0.7 on 30 cases and gets recalibrated quarterly; below that threshold, do not trust the number.

Artificial Intelligence

AI-assisted code review: an honest adoption story

Two years running AI-assisted code review in a real team leave a clear balance: AI catches mechanical oversights well and writes useful pull-request summaries, but it struggles with architectural judgment and produces many false positives on subtle bugs. The single decision that helped the most was not blocking merges on its automated comments.