Your on-call team Is burning out: here's how to see it coming
Introducing On-Call Health, an open-source way of detecting responder overload.
Gandhi Mathi Nathan Kumar
Principal Incident Commander
The first 15 minutes of an incident decide more than the next two hours. Gandhi Kumar explains why mitigation beats diagnosis, silence costs trust, and every extra click matters.

Introducing On-Call Health, an open-source way of detecting responder overload.
Cliff Snyder
Senior SRE at Multimedia
How do you move incident response from 600 specialists to 6,000 engineers without lowering the bar? Cliff Snyder shares the 18-month rollout, its compromises, and the metric that exposed what MTTR missed.
Ganesh Datta
Co-Founder & CTO at Cortex
AI didn't fix your engineering bottlenecks. It put them on fast-forward. Ganesh Datta explains why platform and SRE teams share the same product problem—and why ownership comes first.
Dana Lawson
CTO at Netlify
SREs spent a decade imagining self-healing systems. Now that AI agents are making them possible, many don't want to let go. Dana Lawson gets honest about fear, identity, and learning to trust the tools.


Responding to emergencies is a draining job. No matter how much tools evolve, incident management maturity hinges on humans.
Will Wilson
CEO & Co-founder, Antithesis
The bugs that cause real incidents are usually the ones nobody thought to test. Will Wilson explains how deterministic simulation searches for them—and why AI-generated code raises the stakes.
Stephen Townshend
SRE Team Lead
Burnout didn't arrive with a dramatic collapse. It started with a spot in Stephen Townshend's vision, sleepless nights, and a job with pressure but no control. His recovery is a warning—and a blueprint.
Swizec Teller
Bestselling Author
AI can produce the first 90% of an application in minutes. Swizec Teller is interested in the other 190%: the judgment, ownership, and production scars that turn code into a service people trust.

Learn how on-call policies impact sleep, stress, and burnout, and how fair scheduling, recovery time, and alert control protect long-term support wellbeing.
Dileshni Jayasinghe
VP of Technology at commonsku
commonsku had strong uptime—and no formal incident process. Dileshni Jayasinghe started by giving non-engineers real operational access, then built the guardrails that made shared ownership safe.
Page 2 of 9 · 99 entries