Skip to content

CUSTOMER STORIES

What changed, by how much, and what did not move.

Three teams at different sizes and on different operating models. Each one states the problem it started with, the single decision that changed things, and the numbers ninety days later — including the ones that stayed flat.

HOW TO READ THIS PAGE

Three things worth knowing before the numbers.

The numbers are ours, measured our way
Each figure below compares the ninety days before onboarding with the ninety days after. They are not audited, and no team here was paid to take part.
The flat lines are in the table too
Every story names something that did not improve. A page with three wins and no cost is a brochure, and you would be right not to trust it.
Scale matters more than the percentage
A nine-engineer team and a forty-engineer team fail differently. The team size and the number of services in scope sit at the top of each story for that reason.

Payments infrastructure

Northbeam

40 engineers · 6 services in scope

3h 40m11 min

Median time to first human seeing a checkout failure

Checkout failures surfaced through support tickets, usually a few hours after the first customer hit them. Nobody was ignoring the error tracker — there were simply four thousand open issues in it, and no way to tell which one was costing money right now.

Payment and authentication paths were scored highest during scoping, so anything touching them clears the notification threshold on its own. Everything below high risk stops at the workspace and waits for the weekly review.

What did not improveDeploy frequency did not change, and we would not claim it should have. This shortens the distance between a failure and the person who can judge it — it does not ship the fix for you.

Logistics platform

Halcyon

12 engineers · 3 services in scope

172

Out-of-hours pages per month

A two-person on-call rotation was being paged for everything a monitor could detect, including the same expired-token error four nights in a row. The team had started muting channels, which is the point at which alerting stops working entirely.

Risk levels allowed to interrupt a person were cut to critical only during out-of-hours, with everything else batched into a morning brief. Repeat incidents are grouped, so the fourth night of the same failure is one line in a digest rather than a fourth page.

What did not improveThe first three weeks were noisier, not quieter. Tuning a threshold means finding out where it is wrong, and that only happens by letting some things through.

Developer tooling

Cadence

9 engineers · one product

0

Incident payloads sent to a third-party provider

Customer code was in scope for analysis, so nothing could leave the network — which ruled out every hosted incident tool the team had looked at. The alternative was continuing to read stack traces by hand.

The workspace runs inside their own infrastructure against a local model, so no incident context reaches an external provider. They operate it; we are not in the path at all.

What did not improveRunning a local model costs hardware and someone to maintain it. For this team that was the cheaper trade; for most teams the managed service is less work.

YOUR NUMBERS WILL BE DIFFERENT

The first sync usually finds something failing quietly.

Start with the application that breaks most often. A scoping call is enough to say what we would watch and what we would leave alone.

  • One application is enough to find out
  • No agent to install and no change to your application
  • Sensitive fields are redacted before anything is analysed
Talk to us See what each package includes

A scoping call, not a sales sequence.