How Prolific outgrew incident.io and built an AI-assisted reliability operating model on Rootly.

64%
reduction in time to resolve
100%
completion for retrospectives
86%
reduction in repeat incidents

“Rootly has been more than a tooling upgrade. It's a new operating model for how Prolific navigates failure; faster, with less friction, and with reliability treated as a shared responsibility rather than an afterthought.”

Hannah HammondsService Delivery Lead

Prolific is a human data and research platform that connects AI developers, academic researchers, and businesses with a diverse, vetted pool of real people for online studies, surveys, and AI training tasks. Prolific’s mission is “To revolutionize how the world learns about people, so people can revolutionize the world.”

Founded: 2014 in London, United Kingdom

Size: ~400 employees

When Hannah Hammonds evaluated incident tools for Prolific, she didn’t ask the usual question. Not “does this work today,” but “will this still work in three to five years, when we’re routing more of our process through AI.” incident.io had gotten Prolific started, but as she tried to mature the practice it became a ceiling; limited flexibility, required workarounds, and not built for the AI-assisted future she was designing toward. So she didn’t shop for another pager. She chose Rootly as the backbone for Prolific’s reliability and AI strategy, displaced incident.io, rebuilt the entire incident lifecycle, and ended up helping co-design the AI SRE her team now runs on.

“incident.io got us started, but the moment I tried to mature us past survival mode, it became a ceiling.” —Hannah Hammonds, Service Delivery Lead

When “good” still felt like survival

Before Rootly, a great month at Prolific was defined by the basics holding; incidents coordinated smoothly, people feeling safe enough to raise an incident when something broke, communications that didn’t slip, retrospectives that actually happened, and a team that didn’t burn out on the same problems repeating. Even hitting that bar felt like pushing uphill. The process leaned heavily on individuals doing the right thing under pressure, and there wasn’t enough structure left over to move from survival mode toward a genuine gold-standard incident culture.

Prolific was on incident.io at the time. It had helped them get started, but as Hannah tried to mature their approach, it began to feel like a ceiling. Different teams hit different pain points, workflows felt limited, and day-to-day use wasn’t intuitive for everyone, but the common thread was a lack of flexibility and workarounds piling up. It was difficult to customize the system around how Prolific actually worked, and it didn’t offer the depth Hannah needed for the future she had in mind: incident response that was not only well-orchestrated, but AI-assisted and reliability-driven by design.

The pain got concrete and hard to ignore. Recurring incidents without a clear picture of why, missing context about which services were impacted and how often, and heavy manual effort just to keep the process running. Hannah stepped back and reframed the problem. She wasn’t choosing another pager or incident dashboard. She was designing the backbone for Prolific’s reliability and AI-assisted incident strategy.

“Rootly is the only platform that’s easy to scale, has the flexibility to meet different requirements for different teams, and the AI-assisted capability for the future of reliability.” —Hannah Hammonds, Service Delivery Lead

Choosing for where reliability is going, not just where it is

When Hannah evaluated options, including PagerDuty, she ran the decision on a three-to-five-year horizon. The question wasn’t only does this work today, it was will this still work when we’re routing more of the process through automation and AI. She also knew any tool would fail culturally if it were seen as just for engineering or just for service management, so it had to be something the whole company, engineers, managers, and stakeholders alike, could see value in.

Rootly’s core product gave her what she had been missing; highly configurable workflows that were easy to manage, deep Slack integration, strong timelines and audit trails for compliance, and robust communications that elevated how Prolific communicated both internally and externally. Just as importantly, Rootly had built, and was continuing to build, the AI-powered capabilities she had been imagining. Leading Prolific to displace incident.io in favor of Rootly wasn’t a lateral move between similar tools. It was a strategic decision to partner with a platform, and a team, that matched her ambition.

“We chose Rootly as the backbone for where Prolific’s reliability and AI strategy needed to be in three to five years. The question wasn’t does this work today, it’s will it still work when we’re routing more of the process through AI.” —Hannah Hammonds, Service Delivery Lead

What changed

A repeatable lifecycle, not a hero effort

The implementation was intentionally full-stack. Hannah rebuilt Prolific’s incident lifecycle end to end: an incident is raised, by a system or a person, and Rootly spins up the right channel, alerts the right people, starts building the timeline, automates the root cause analysis, and scaffolds PIRs and follow-ups automatically. Crucially, training wasn’t kept inside service delivery. She brought engineering managers, business owners, and stakeholders in early so Rootly worked across the whole organization.

The payoff was consistency. Incidents no longer depend on who happens to be on call, they follow a clear, repeatable pattern. Alerts reach the right responders faster, time to acknowledge and resolve dropped, automated communications reduce the risk of missed steps under stress, and standardized workflows make incidents predictable instead of ad hoc, which makes audits easier and gives leadership a clearer view.

On-call that nobody faces alone

Hannah designed a primary and secondary escalation structure so no one is truly alone when something breaks. Rootly’s flexible escalation lets Prolific bring in extra support with clear context, whether pairing on a rollback or involving another team, and communicate why the escalation is happening. Out-of-hours incidents are streamlined; alerts automatically create incidents, notify the right people, and ensure the critical steps are followed every time. Paired with an observability culture that treats false positives as opportunities to improve alerts, this shifted the mindset from “who’s on the hook” to “how do we make the system better.”

Root cause that’s real, not a checkbox

A core part of Hannah’s philosophy is that root cause analysis should be real, not a box-ticking exercise. Prolific doesn’t stop at broad labels like “deployment” or “human error.” Once they know the category, they run a 5 Whys analysis and treat post-incident reviews as collaborative sessions anchored on how they could have detected the issue earlier, reduced its impact, and which improvements deliver the best value for the effort.

To make that sustainable, Hannah designed a workflow around Google Meet transcripts. When a PIR starts, the transcript is captured and automatically structured around their key focus questions, then engineers refine it. That improves the accuracy of technical details, reduces the manual burden on the PIR lead, and makes lessons learned faster to distribute and more grounded in what actually happened.

AI SRE that reinforces the process

This is the part Hannah had been building toward, and the partnership deepened over six months through the development of Rootly’s AI SRE, which she and Prolific helped co-design. It connects to Prolific’s monitoring stack, analyzes relevant logs, alerts, and PRs, and brings probable root cause directly into the incident channel.

In practice it has shaved valuable minutes off the early triage window, cut the time from “something’s wrong” to “we know where to look,” and reduced dead-end investigations, even drafting rollback pull requests when deployment issues are identified. The result is quicker acknowledgements, more accurate first responses, and less variation between incidents.

It is especially powerful out-of-hours; when fewer people are online and context is thinnest, it gives on-call engineers a structured set of steps so they aren’t starting from a blank slate at 3am, which means faster stabilization and fewer “wake the entire team” moments. Stakeholders can even ask @Rootly what’s happening to get a clear, current answer without disrupting responders. The point, for Hannah, is that AI SRE doesn’t replace the process they designed, it reinforces it, so the lifecycle runs the same way every time regardless of who is on call or how tired they are.

A new operating model, and a culture

Hannah is candid about the biggest gotcha she sees other teams fall into; assuming a tool alone will fix a messy process. Without clear workflows, ownership, and expectations, even the best platform produces chaotic incidents. Her approach was to lead with understanding, listen to teams’ priorities and pain points, tie everything back to delivering a reliable service to customers, and position herself as an enabler rather than a gatekeeper. Clear communication, early involvement, and quick visible wins built trust that accumulated into lasting cultural change. Today Prolific has an incident-first culture and a deeply integrated Rootly deployment with AI-assisted capability the team helped co-design.

Prolific’s lightning round, scored 1–5

  • Pager fatigue lower: 5
  • Time to first status update improved: 4
  • Retro completion improved: 5
  • Stakeholder comms are consistent: 5
  • Engineering morale improved: 5
  • Response and resolution times decreased: 5
  • Customers are happier: 4

The tool that gets a team off zero is rarely the one that takes it to a gold standard, and almost never the one built for where reliability is heading next. Prolific’s answer was to stop treating the decision as a pager swap and start treating it as a design problem; what backbone does an incident-first, AI-assisted reliability practice actually need. Rootly was that backbone, configurable enough to match how Prolific works, rigorous enough and forward enough that Prolific helped co-design the AI SRE it now runs on. If incident.io got you started and you can feel the ceiling, the question worth asking is the one Hannah asked. Not what works today, but what will still work when more of your process runs through AI.