# The role of automation in incident response

URL: https://rootly.com/incident-response/automation
Pillar: Incident Response
Bundle: https://rootly.com/llms/incident-response.txt

> Automation runs the incident response steps a team can define in advance, so responders spend the incident on diagnosis and decisions. What to automate at each stage, what to keep human, and where to start.

Every incident starts with the same chores: someone declares it, opens a channel, pages the right people, starts a bridge and tells the stakeholders. None of that finds the cause, and all of it competes for attention with finding the cause. Automation is how a team takes those chores off the responders. Rootly runs them as workflows, so this page uses Rootly as the worked example.

## What is the role of automation in incident response?

Automation runs the steps of incident response that a team can define in advance: declaring the incident, opening the channel, paging and assigning roles, posting reminders to update stakeholders, creating tickets and starting the retrospective. That leaves the responders' attention for the parts that need judgment, which are working out what is wrong and deciding what to change. In Rootly, each of those steps is a workflow with a trigger event, run conditions and actions, and it runs automatically or from a Slack command.

Automation earns its place three ways:

- **Consistency.** The same steps run on every incident of a given severity or service, including the one that starts at 3 a.m. with a responder who has never led one.
- **Speed of setup.** The channel, the bridge and the pages exist before anyone has finished reading the alert.
- **A complete record.** Every automated step lands on the incident timeline, so the retrospective starts from the recorded sequence of events.

## What should you automate at each stage of an incident?

Automate the steps with one right answer, and keep a person on the steps that depend on reading the situation. Rootly workflows cover the automated half of each stage:

- **Detection.** Automate: declare an incident when an alert your team has already vetted as needing immediate human action arrives, such as a customer-facing SLO breach, and page the on-call responder for the affected service. Keep human: deciding severity when the signals conflict.
- **Mobilization.** Automate: create the incident channel in Slack or Microsoft Teams, start a Zoom or Google Meet bridge for high severity, assign roles, pull in responders. Keep human: deciding who else the incident needs.
- **Investigation.** Automate: gather related alerts, recent deploys and past incidents into the channel, and start an AI SRE investigation. Keep human: choosing which explanation to act on.
- **Mitigation.** Automate: run runbook steps with a known, safe outcome. Keep human: approving changes with real blast radius.
- **Communication.** Automate: post periodic reminders to update the status page; notify legal, security or support when high-impact incidents occur. Keep human: the wording of customer-facing messages.
- **Follow-up.** Automate: create Jira or Linear tickets for the affected team; create the retrospective document from a template. Keep human: the analysis of contributing factors and what to change.

The "keep human" half is where AI helps without taking the decision. Rootly AI SRE investigates and reports the most likely cause with its evidence, and Rootly AI drafts status summaries, so the person deciding starts from a draft. [What AI incident management is and how it works](https://rootly.com/ai-sre-guide/ai-incident-management) covers that side.

## What should stay manual?

Keep a person on any step where a wrong answer is expensive and the right answer depends on context. Three are common:

1. **Production changes with a large blast radius.** A rollback that is safe for one service can take down three others. Automate the diagnostics around it and let a responder approve the change.
2. **Severity when signals disagree.** An error spike with no customer reports might be a sev-3 or the start of a sev-1. Automate the declaration at a default severity and let the incident commander adjust it.
3. **External wording.** Automate the reminder to update the status page and draft the text, and have a person approve what customers read.

## How do you start automating incident response?

Start with the steps you already do on every incident, then add conditions. Rootly workflows are built in the same order:

1. **Write down what happens on every incident.** The list is usually short: declare, open a channel, page, start a bridge, tell stakeholders, open a retrospective.
2. **Automate declaration and the channel first.** They run on every incident and have one right answer.
3. **Add conditions by severity and service.** A sev-1 on the payments service may need legal and support notified and a bridge started; a sev-4 needs neither.
4. **Automate communication reminders.** A repeating reminder to update the status page removes the most common gap in stakeholder updates.
5. **Automate follow-up.** Create tickets and the retrospective document when the incident resolves, so learning starts before people move on.
6. **Review the automation after each retrospective.** Every manual step that happened twice is a candidate for the next workflow.

Keep [runbooks](https://rootly.com/incident-response/runbooks) alongside the workflows: a workflow can post the right runbook into the channel, and the runbook holds the steps a person still runs.

## Is incident response automation the same as SOAR?

They share a principle and differ in scope. SOAR (security orchestration, automation and response) platforms automate security incident response: triaging security alerts, enriching them with threat intelligence and running containment steps. This page covers incident response for software reliability, where the incident is an outage, a degradation or a failed deploy. In both, automation takes the defined steps off the responders so they can spend their attention on judgment. Rootly is built for the reliability side.

## Frequently asked questions

### What is the role of automation in incident response?

Automation runs the incident response steps a team can define in advance, such as declaring the incident, opening the channel, paging responders, posting stakeholder reminders, creating tickets and starting the retrospective. Responders then spend the incident on diagnosis and decisions. Rootly runs these steps as workflows, triggered automatically by incident and alert events or manually from Slack.

### What parts of incident response should be automated?

Automate the steps with one right answer: declaration, channel and bridge creation, paging, role assignment, stakeholder notifications, ticket creation and the retrospective document. Keep a person on severity calls when signals conflict, production changes with a large blast radius, and customer-facing wording. Rootly workflows handle the first group, and Rootly AI drafts for the second so a person decides from a starting point.

### How does automation reduce MTTR?

It removes the setup time at the start of every incident and the coordination steps during it, so responders start diagnosing sooner and fewer steps get missed. It shortens MTTR most on the incidents where setup was slow or inconsistent. Rootly records every automated step on the incident timeline, which lets the team see where time went in the retrospective.

### Can incident response be fully automated?

The defined steps can be, and the decisions shouldn't be. Declaring, paging, notifying and ticketing run well without a person. Choosing which explanation to act on and approving a risky production change still need a responder with context. Rootly automates the first group with workflows and supports the second with AI SRE investigations that show the evidence behind each conclusion.
