SRE is a techno-social craft
Tony Holmes was a sysadmin before SRE had a name. Now leading SRE at Affirm, he reflects on what the craft retained, how severity guidelines fail, and why mentorship is leadership infrastructure.

Tony Holmes
Head of SRE
- Ex-Apple, YouTube & Netflix
- Tech mentor
- PC Gamer
- 30 year tech vet
SRE is a techno-social craft
On this page
Key Topics Discussed
- What modern SRE inherited from traditional systems administration
- The decisions and responsibilities of leading an SRE organization
- Designing severity guidelines teams can apply under pressure
- Mentorship as a core responsibility for senior reliability leaders
- Life beyond the incident queue
Where to Find Tony Holmes
- Watch: Watch the conversation
Transcript
You got your start as a sysadmin long before the term āSREā had even entered the lexicon. How do you view the difference between sysadmin and SRE work today? Is SRE an evolution of the sysadmin role, or a totally different craft?
SRE is a techno-social craft. You have to take the social side of the equation equally with the technological side, because youāre dealing with people using systems or trying to interact or bring systems back up into good state, so you need to consider the human conditions in that moment.
I think of SRE as a venn diagram. You have your administrators of network systems, storage administrators, developers, devops engineers, CI/CD engineers, and so on. Approximately 10% of them will approach it from a reliability-first perspective. This means asking, first and foremost, āHow will this thing fail? And, in the face of that failure, how will we respond?ā When this mindset is the first reflex in a discipline, thatās when you get a reliability engineer. I have seen so many companies get it wrong. They had sysadmins and said āOkay, youāre now SREs.ā No. The skill set that you need is that curiosity and holistic view of the system, in addition to the ability to debug and understand the technology itself. In SRE, youāre looking beyond the systems in isolation, and considering the whole experience that your systems create for users.
Your role at Affirm is the Head of SRE. What does that look like for you day-to-day? What type of decisions are you making and what actually consumes a lot of your time leading an SRE organization?
Yeah definitely. So right now, itās kind of a dual role. Iām leading the team while also redefining the function as a whole, from how it was run under previous leadership. Iām making little nudges, tweaks and optimizations here and there but thinking about the big picture defining the program goals and methodologies. Right now, itās working closely with my team to recognize their pain points and address them to support their ability to execute on their mission. One of the things weāre working on now is our incident management flow. Weāve identified opportunities to reduce what we call āself-inflicted painā ā we identified that 73% of our incidents were being caused by human changes. Weāre also redefining our severity guidelines so they work inline with our SLOs. So, my goalāto sum all that stuff upāis Iām trying to create and execute on a framework that factors in the human side of the equation, making things clearer, more consistent, easier to reason about.
You mentioned rethinking severity levels - letās dive into that a bit more. Tell me about what makes severity guidelines effective.
Google does it really well. For anyone who hasnāt read the Google SRE book, their chapters on defining SLOs and multi-window alerting is a great place to start. You want to base your alerting on burnāerror budget burn. If you are delivering three nines of reliability, that means you have an error budget of 0.1% over whatever period. So for us, we have monthly SLOs, three nines will be one of them. Some services are tighter, some are more loose, but Iāll use that as an example. We want to alert urgently if weāre burning the error budget very quickly. So if thereās a major incident, maybe a system like DNS has gone out, which will impact a lot of things, which means youāre burning a lot of SLOs broadly everywhere. Youāll want your alerting to indicate that. So weāll actually base severity on our error rates against our normal volume of transactions. A SEV0 essentially means that over a very short period of time, youāre going to burn a very significant chunk of your SLO. As an example, letās say SEV0 means you are going to burn 10% of your error budget in six hours. Thatās a lot. So you want your alerts aligned around that, and then you set the various levels. So the SEV0s will be your very, very sharp ones. Your next one will be the SEV1s, which is probably going to be 1% in the same period or 10% over a longer period of time. Google has a pattern of six to one ā where the height versus the width, it tends to be a six to one ratio for all the definitions. So a SEV1 is six times more impactful than a SEV0 in terms of the window that they look over.
As a senior leader in SRE, how do you approach mentorship? What does being a good mentor mean to you?
First, there are a lot of different types of mentorship. I have a mentee from Toronto that Iāve known for 14 years. Heās a close friend and a mentor around life goals, aspirations, direction, and we meet once a month and we have these really intense conversations and we meander through our thoughts, our focuses, and our curiosity.
At Affirm, I believe that everyone in our team should be mentoring somebody if in some capacity, with the exception of the juniors whose entire reason to be is to learn inside the team. You have the technical role-based mentors. I will typically mentor the most senior engineers in my team. They will then mentor the next layer down, so on and so forth, and thereāll be some skips and stuff going on in there to give perspective.
So thereās technical mentorship, which is easy. This is a skill, hereās how I recommend you approach it, hereās how you go and learn it. āSoftā skills, thatās the harder one. There are a lot of people that have a very difficult time going from a technical to a soft influence based skillset. So the ones that can do that well, thatās a superpower. Thatās when Iāll say āI want you to go talk to this person because they are really good at influence without authority. I want you to go and spend six months with them, whatever cadence works, 30 or 60, 30, 60 minutes a week, and I want you to make it very clear this is what youāre looking for and you want guidance on how you can grow that skill.
I also recommend people have mentorship from outside of their immediate team, like meeting with another high level IC from another team so you can understand how they perceive your team. And then you want to go to a product team. How do they view the role of our platform, our team? The mentorships should be with a purpose, and some of them, once you get the skill, they should end and then you will find another one for the next skill you need to build.
Affirm is a huge player in the ecommerce space. Are you a big online shopper? Whatās the best purchase you made online recently?
My wife is the big online shopper in our house, but Iāve done my fair share and Iām actually staring at mine right now. It is a wonderful large curve monitor. Affirm has a āwalletā program that allows you to buy equipment for your home office, since we primarily work remotely. So I upgraded my work monitor, which is fantastic, because I also get to use it for gaming. I got it two months ago and Iām absolutely loving it.




