Could 2027 be the year you turn off human-in-the-loop remediation?

Installing and configuring a reliability platform is already code. The agents already have the context. The only step still gated on a human is the change itself — here is what would have to be true to ungate it.

Every part of running an incident has been automated except one.

You can describe your reliability platform in Terraform and apply it alongside the infrastructure it watches. Schedules, escalation policies, severities, workflows, status pages—all of it as code, reviewed in a pull request, rolled back like anything else. Agents can read your service catalog, on-call schedule, and deployment history; query telemetry sources; and understand years of incidents, including your retrospectives. A service graph and memory with custom notes are part of the model that learns and understands your environment. When something breaks, agents correlate telemetry with recent changes, incident history, and memory to produce a ranked set of probable causes, each with a confidence score and the reasoning behind it.

Agents then determine the next course of action and a suggested fix. Then they stop and ask a person to approve.

That pause is a deliberate product decision, not a missing feature. Rootly does not auto-remediate without human sign-off. The question worth asking, with the end of 2026 in sight, is: should 2027 be the year a team turns it off on purpose?

Your on-call config is already code

Start with what is settled, because it is more than people assume.

A reliability platform is infrastructure, and you configure infrastructure through a provider, not a browser. A rotation becomes a diff. Onboarding a service becomes a module. The question of who changed the escalation policy before an outage gets answered from version control instead of memory. Add an API and an MCP server, and the platform is addressable by the same agents that manage everything else in your stack.

This is the part of the autonomy story that is finished. Nothing about installing, configuring, or changing a reliability platform requires a human hand on a mouse. If on-call configuration were the obstacle to autonomous operations, we would already be there.

Agents already assemble what a senior engineer builds by hand

The second thing people underestimate is context.

An agent investigating an incident is not reasoning from the alert alone. It has the service that fired, who owns it, what shipped in the last hour, day, week, month, which of those changes touched the failing path, what the error rate looked like before, and whether this shape of failure has happened before and what closed it last time.

That is roughly the full set of inputs a senior engineer assembles in the first fifteen minutes of a serious incident. Assembling it is the part of on-call that is genuinely mechanical, and handing it to software is straightforwardly good. It is also most of the job, by time.

The only gated step is the change itself

What remains is the change itself. Restart the pods, roll back the deploy, fail over the region, drop the feature flag, merge the PR with new code…

The reason to gate it is not that a model cannot propose the right action. It can, and it can explain why. The reason is that a model has no reliable sense of when it is out of its depth. Presented with a failure mode it has never seen, an agent with write access does not hesitate — it acts on its best available guess with the same confidence it brings to a case it has seen a thousand times. On a novel edge case, that is how a contained incident becomes a compound one.

The public evidence supports the caution. On OpenRCA, a benchmark for autonomous cloud root-cause analysis, leading agents score between 3.9% and 12.5% on fully correct end-to-end diagnosis. Those numbers describe end-to-end autonomous diagnosis without a human checkpoint, which is precisely the configuration anyone turning off human-in-the-loop would be running.

Our own benchmark says models aren’t there yet

Arguing about whether models are “good enough” without an instrument is a waste of everyone’s time, so we built one.

SRE-skills-bench is a benchmark from Rootly AI Labs that evaluates models on some of the work SREs actually do: triaging an incident, reading logs, suggesting a mitigation. It was featured at ICML and ACL 2025, and it now runs inside Groq’s OpenBench, so anyone can reproduce the results rather than take a vendor’s word for it.

What it shows is a real trendline. Frontier models keep getting better at SRE-shaped reasoning, and the gaps between them are narrowing and shifting release to release. What it does not show is any model clearing the bar that unattended production changes would require. Measuring the slope is not the same as knowing where it crosses.

Four things have to be true before you turn it off

What would a team need before flipping the switch? Four things, and none of them is a bigger model.

A calibrated sense of not knowing. The single most important capability is abstention: an agent that recognizes an unfamiliar failure and escalates rather than guessing. Confidence scores are a start, but a score is only useful if it is calibrated, meaning that the things labeled 90% are right 90% of the time. That is measurable, and you should measure it on your own incident history before you trust it.

Impact radius as a first-class control. Autonomy is not one switch; it is a dial per action. Restarting a stateless pod in staging and failing over a primary database are not the same decision and should never sit behind the same permission. A credible autonomous mode lets you enumerate exactly which actions an agent may take unattended, on which services, in which environments, within which time windows.

Rollback guarantees that do not depend on the agent. If an autonomous action makes things worse, recovery cannot require the same system that made the bad call to notice and correct it. The undo path has to be external, automatic, and verified before the action is permitted, not after.

Evidence a human can audit after the fact. Every unattended change needs a record showing what the agent saw, what it considered, what it rejected, and why it acted. Without that, the first time an autonomous action coincides with an outage, the team loses confidence permanently and turns it off. Trust is built on reviewable decisions, not good outcomes.

Opsgenie’s April 2027 shutdown forces the question

The date is not arbitrary. Atlassian ends support for Opsgenie on April 5, 2027, after ending new sales on June 4, 2025. Thousands of teams have to move their paging stack before then, which means thousands of teams are running an evaluation they would otherwise have deferred for years.

Re-platforming is the moment teams reconsider defaults. When you are rebuilding rotations and escalation paths anyway, the question of how much a human needs to be in each loop is suddenly live in a way it never is during business as usual. Whatever gets decided during those migrations will set the operating model for the rest of the decade.

The second reason is that the reasoning itself stopped being the bottleneck. A model that can read a stack trace, hold a distributed failure in its head, and argue from a latency graph to a probable cause is no longer remarkable; that capability arrived and kept improving. Judged purely on reasoning, the case for autonomy looks stronger every quarter.

Which is exactly why the OpenRCA result above matters more than it first appears. Those agents failed on architecture, not on reasoning. They were smart enough and still got it wrong, because they were reasoning over whatever they happened to be handed.

That is the gap a context layer closes, and it is the part of the problem that is actually ours to solve, not a frontier model vendor’s. Rootly’s AI has full context across services, team ownership, on-call schedules, incident history, and telemetry before it starts investigating, then runs parallel hypothesis checks across alerts, telemetry, recent deployments, and past incidents to build a ranked, evidence-backed theory—each finding carrying a confidence score and a visible reasoning chain. The reasoning is the model’s. Knowing which service just shipped, who owns it, what broke last time it did this, and what the graph looked like an hour ago is the platform’s.

So the trend line worth watching into 2027 is not model benchmarks. It is whether the context handed to a good model gets complete enough, and structured enough, that its conclusions become reliable rather than merely plausible. That is a build problem with a visible finish line, which is why the question is worth asking now rather than in 2030.

Rootly keeps the human in the loop, for now

Almost every remediation an agent proposes goes to a person by default.

We think that default is right for now, and we think the interesting question is what would change it. Our answer is that it is not a model capability threshold on its own. It is calibration, scoped impact radius and rollback, and auditable evidence; engineering problems, mostly, and tractable ones. Problems we’ve solved but are determined to keep a human in the loop, for now. Our team has all four reasonably running against a narrow class of well-understood, low-impact-radius actions unattended and keeps humans on the rest. We always dog-food our own product experiences.

That is a meaningfully different claim from “nobody should be on-call in 2027.” Nobody should be on-call the way they are in 2026—woken for noise, reassembling context by hand, paging into a tool that is being switched off. That part is already solved. The rest is a question we want people to challenge as we release the trust you need to enable autonomous actions and self-healing.