Can AI Actually Find the Root Cause? Automating Postmortems and RCA
What automation reliably contributes to a post-incident review, what it cannot, and how to run retrospectives that produce fixes instead of documents.

What automation reliably contributes to a post-incident review, what it cannot, and how to run retrospectives that produce fixes instead of documents.


Installing and configuring a reliability platform is already code. The agents already have the context. The only step still gated on a human is the change itself — here is what would have to be true to ungate it.

What AI SRE agents reliably do today, what they still get wrong, and how to evaluate a vendor claim without taking the demo at face value.

We built Doom Agent Arena, an open-source benchmark where AI agents battle in Doom via MCP. Here's what it taught us about AI-assisted incident response.

Brent Chapman argues that AI-written post-incident reviews quietly automate away the learning. We build the AI that drafts them.

Anthropic released Claude Sonnet-4.6, and we ran it through SRE-skills-bench the same day. It tests models on the tasks SREs actually do: understanding infrastructure code, reasoning about cloud configurations, and mapping code diffs to real-world pull requests.

AI SRE brings AI to incident response, root cause analysis, and remediation, reducing on-call load and improving reliability outcomes for teams.

A shift just happened in SRE AI performance. Gemini 3 Pro didn’t just edge out OpenAI’s models, it beat them across every SRE task we threw at it. The landscape is changing faster than anyone expected.

5 takeaways from Atlanta on AI, Kubernetes, and reliability

Why 50% of companies don't monitor ML and how it’s reshaping our understanding of reliability.

The new edition of our benchmark features Terraform tasks across AWS, GPC, and Azure, plus incorporates a new dimension: prompt-optimization.

The panel warned: the opportunity is massive, but without observability, security, and strategy, the regrets will be real.

Making LLM evaluations reproducible for real-world SRE workflows
Page 1 of 3 · 27 articles