Back to customers

How UNiDAYS replaced a decade of PagerDuty in two weeks and opened incident response to the wider business.

UNiDAYS

x

rootly-logo
“We moved ten years of incident tooling to Rootly in two weeks, and now the whole company can take part.”
Saag PatelUNiDAYS

Saag Patel

,

Head of Platform

UNiDAYS is the world's largest student affinity network, connecting more than 29 million verified student members across 115 markets to over 800 brands, on a mission to support, enable and inspire young people to be their best selves. Alongside the marketplace, it now sells student verification as a standalone product to brands that need eligibility confirmed rather than a discount offer.

Founded: 2011, Nottingham, United Kingdom

Size: ~300

Rootly’s Impact

Hours → minutes

to produce an incident timeline

Engineering-only → business-wide

operations

Hundreds

of engineering hours saved

For about ten years, UNiDAYS paid PagerDuty to do exactly one thing: page people. Not because that was the plan, but because the seat count on their tier meant only a small group could ever get in. All the incident-response capability sat behind that ceiling, so the platform team improvised around it, pushing retrospectives into Confluence where more colleagues had access, rebuilding incident timelines by hand, and catching up the stakeholders who got missed. Then renewal arrived with another increase and nothing new to show for a decade of loyalty, and the team decided to scan the market instead. Less than two weeks later they were live on Rootly, and the timelines that used to take hours were taking minutes. Saag Patel, who leads platform, has a sharper way of describing what they left behind. 

Ten years of best endeavors

UNiDAYS runs against committed SLAs and ships fast with small teams, working toward what Saag calls a state of equilibrium: let products evolve quickly without reliability becoming the dominant source of noise. That takes a real incident process, and for most of a decade the tool could not supply one. The seat count was a budget decision as much as a licensing one, and its effect was that PagerDuty stayed what its name suggests. It paged. The team did not meaningfully evolve their usage until the last couple of years, and with little Slack integration and thin third-party integrations, everything around the page stayed manual. Saag describes the whole model as best endeavors: get a few people in a room and resolve it as fast as you can.

The friction compounded from there. When the team asked how to automate further, the answer was to move up a tier, which the budget choice did not agree with. They had paid for live call routing and were never able to switch it on. Support was attentive at renewal time and, in Saag's assessment, a ticket in a queue or quietly closed for the other eleven months. There was no roadmap visibility either, so nothing worth waiting for. When the renewal came around with the usual five to seven percent and no feature development behind it, the choice was the one Saag puts plainly: break the cycle now, or iterate for another year and add another increment.

"Rootly gave us the flexibility for every edge case, workflow, anything we could think of, and it’s easy to work with; the UI is intuitive." -Saag Patel, Head of Platform

A decade of configuration, moved in two weeks

The migration is the part other platform teams will want to copy, and it was deliberately unglamorous. Five steps, roughly in this order:

  1. Start from code you already have. UNiDAYS already managed PagerDuty in Terraform, so Usama's first move was to translate that into Terraform for Rootly. That one decision did most of the work: the team knew exactly what existed and could reproduce it rather than rebuild it.
  2. Port the alerting model as-is. PagerDuty's service-based model came across unchanged, so an alert still hit a service and still called the same escalation policy. Continuity over reinvention, to keep the cutover speedy and boring on purpose.
  3. Bring the alert sources across and test them. SNS, CloudWatch, and New Relic were plugged in, and the email-to-incident workflow was rebuilt and tested with the team before go-live.
  4. Run one week of proof of concept, then one week of big bang cut over. Two weeks, end to end, with nothing going off that shouldn't have.
  5. Lean on the vendor during the window. Whenever the team flagged something, they say Rootly sorted it, which for a team used to a support queue was its own kind of proof.

Value arrived inside that same fortnight, in two forms. Financially it was immediate: as Saag puts it, the price of one tool against the price of another, and the delta is what they bank. Operationally, it landed on the first real incident, where historically some stakeholders would have been missed and caught up by hand, and instead the right people were simply there.

“Setting up Rootly was smooth. The support was very clear and transparent, and also quick. It was as smooth as it can be." -Usama Blavins, Senior Platform Engineer

What changed

The seat ceiling was the real problem, and it's gone

This is the change that matters most, because it explains all the others. On PagerDuty, capability existed that UNiDAYS could not distribute, so the organization worked around the tool instead of in it. On Rootly, participation is not rationed by a seat count, which turns incident response from an engineering-only activity into something the wider business can join. That shift is already underway on two fronts. Operations functions can now see an incident and join the Slack channel themselves rather than waiting to be pulled in, and the team is rolling the process out to further departments with what Usama describes as much stronger buy-in. Internally, it also changes who owns incident management at all: platform has historically been the gatekeeper, and it is now moving that responsibility out to individual engineering teams, which is only realistic once access is not rationed. 

Timelines that assemble themselves: hours to minutes

Saag's own description of the saving is hours to minutes, and it is the clearest operational win because it recurs on every incident. In the PagerDuty era someone had to populate the incident timeline by hand, genuinely time consuming on a long incident. Now the team pins messages as they go and the timeline is produced for them, ready to feed Rootly’s AI retrospectives. It is, in his words, a subtle piece of value, and anyone who has assembled a timeline by hand will recognize how much of it there is.

Out of the box instead of bolted on

Where the old setup needed heavy custom automation just to reach a baseline, the equivalent now runs by default: the right Slack members are brought into the incident, the surrounding setup happens automatically, and things are tidied up on close. The email-to-incident path is a simple forwarding rule, and crucially the team can watch an alert land and see the workflow fire, where previously the same integration was opaque and would break when keys changed. Templated workflows, escalation policies, schedules, and alert routes are all managed in the same Terraform-first way the team already worked.

Proof points (Company’s internal read-out)

  • 5/5: Customers are happier
  • 5/5: Time to first status update improved
  • 5/5: On-call handoffs are clearer and cleaner
  • 4/5: Stakeholder communications are more consistent
  • 5/5: Engineering morale improved
"There is more calm in the business now, more visibility and transparency. We have a faster path to resolution and a better developer experience." -Saag Patel, Head of Platform

Plenty of teams stay with a vendor for a decade out of inertia, and pay for the privilege in annual increments. What UNiDAYS discovered is that the cost was never only the invoice. It was a seat ceiling that quietly decided how few people could take part in an incident, and a decade of workarounds built to compensate. Replacing it took two weeks, because the configuration was already code and the migration was designed to be boring. They lifted and shifted first and are now iterating into native retrospectives, AI, and private incidents, which is the point: the ceiling is gone, so there is somewhere to grow. If your incident tooling is licensed in a way that keeps most of your company out of the room, that is the constraint worth pricing, not the renewal.

You and your teams deserve
modern incident management.