Categories

Methodologies

The Site Reliability Workbook: patterns we still use

Seven years on, the Site Reliability Workbook still earns its place in small teams through a handful of patterns: SLOs set slightly below what you already achieve, so the error budget is real and negotiable; a 28 to 30 day rolling window; and blameless postmortems that drop punishment while keeping accountability.

Methodologies

Observability and SLOs: Error Budgets That Get Met

SLOs and error budgets only work when the budget drives real decisions. A feature freeze that triggers on exhaustion, deploy velocity that adjusts to consumption. With two or three well-chosen SLIs, a clear freeze policy, and simple tools like Prometheus with Sloth, a team can sustainably balance velocity and reliability in production.