Blog - Page 21

SRE Complete Resume Writing Guide
Follow these steps to write a great SRE job resume.

An Introduction to Incident Response Roles
Learn about the key roles within an incident response team, as well as optional incident roles you may not have thought about.

What SREs Can Learn from Facebook’s Largest Outage
An SRE’s analysis of the October 2021 Facebook outage.

Google’s State of DevOps 2021 Report: What SREs Need to Know
The four key takeaways for SREs from Google’s State of DevOps 2021 report

The Role of SREs in Observability
Although conversation about observability often ignores SREs, SREs have a central role to play in observability success.

Kubernetes Incident Management Best Practices
In this post, Rajesh Tilwani (Co-Founder of Humalect) covers a variety of strategies for preventing and managing incidents with Kubernetes.

You Do the Math: Reliability Issues Triggered by Math Errors
Even seemingly minor math bugs in software code can have outsize consequences.

Making Your On-call and Incident Management Program Stick
Maintenance of your incident management practice is as important as creation - find out what you can do to keep your engineering organization strong and consistent year over year.

SRE vs. Platform Engineering: The Key Differences, Explained
An overview of the similarities and differences between Site Reliability Engineering and Platform Engineering, including from a career perspective.

How to Improve Upon Google’s Four Golden Signals of Monitoring
The Four Golden Signals of monitoring and observability get a lot of things right. But they could be even better.

Incident Management Goes to the Olympics
A look at outages and disruptions to the IT systems that power the Olympics, from 1996 to today.

SWE Meaning: Software Engineer vs. SRE Explained
SWE stands for software engineer. Here is what the role covers, how it differs from an SRE, and where the two overlap on reliability.
Page 21 of 23 · 272 articles




