← Lehua Gray Case study 01 · SRE Agent

PagerDuty · Design owner & AI pillar design lead

When the stakes are high, design for context over cleverness

PagerDuty's flagship AI and its biggest strategic bet: an agent that investigates production incidents on its own, for engineers who have to trust it in the middle of an outage. In a high stakes situation, getting the right context into the agent was the key to success. 

An investigation running start to finish, with the agent's reasoning visible as it goes.
My role
Design lead for the AI business unit: 7 teams and 5 PMs, with two senior designers under my mentorship. SRE Agent was mine directly, coordinating 4 delivery teams.
Research
30+ discovery interviews, moderated testing, and conversational analysis of every real customer conversation with the agent.
Users
Incident responders and NOC teams at enterprises like Checkout.com, NVIDIA, Salesforce, and New York Life.
Result
Highest engagement of any agent PagerDuty shipped. 25% fewer responders needed on major incidents at Checkout.com.

01 — The problem

The model knew everything about systems in general and nothing about yours.

The people who need an SRE agent most are NOC teams: junior responders working partially- or well-known incidents, who need guidance more than they need a research partner. Half the industry was shipping a chat box and calling it an agent — plausible prose about a system it had never seen. This was also the most complex agent PagerDuty had shipped, on a completely new surface.

Everyone was using the same clever frontier models, but none of our competitors had the right context. Our customers wanted an agent that knew what their systems actually look like, what broke there last time, and what this responder should do next. Nobody in the market at the time had solved this.

  • Net new, with no data and nothing to iterate against. Research was the only way to get it right the first time.
  • We needed to ship fast with deliberately scaled-back capabilities, so the design had to find the user and the use case that matched what the agent could actually do.
  • We were shipping against a company-wide OKR, on a company-wide clock.

02 — RESEARCH

New research methods for a new kind of product.

Agents don't fit the research playbook we had, so I brought in methods that did, at every step of the process. That's what let me design improvements before customers thought to ask for them. More than once, a request came up on a feedback call that I had already designed and gotten onto the roadmap.

Discovery & testing

Before I started designing I did 30+ discovery calls with incident responders, managers, and NOC teams. Once flows took shape, I switched to moderated testing in prototypes. What they turned up about who would actually reach for the agent contradicted what the team assumed.

Conversation analysis

After early access, I added a method the org didn't have. I reviewed real conversations in Arize and built a taxonomy of what the agent was being asked, where it broke down, and exactly what was missing. Then I automated it with LLM-as-a-judge tests built on my analysis framework, so the insights kept flowing without me. Notes Analysis came out of this work and became one of the most used features in the product.

Discovery pipeline

I built a pipeline in Claude that mined the customer calls CS was already recording, pulling insights into the synthesis from calls that were already happening. I packaged it as a shareable skill, and design teams across the company now use it to process their own research.

The evaluation pipeline: golden test questions run against the agent version under test and against an expected-response generator, with both outputs scored by an LLM-as-a-judge into an evaluation results dashboard and historical results
Golden test questions run against each agent version and scored by an LLM-as-a-judge, so regressions surfaced without anyone reading transcripts by hand.

03 — DESIGN VISION

Every decision built the right context.

Each one with the alternative I turned down and the evidence that settled it.

Decision 01

Ship it in the right context: web first, not Slack

The assumption was Slack — that's where responders live. Research said otherwise: the NOC users who most needed guidance worked in a console all day, and shipping to the web meant a feedback loop measured in days instead of waiting out a long Slack adoption cycle. I fought for it, and it worked. We had full control over the UI in web and it is purpose-designed for our agents, so teams that later adopted Slack still tell us they prefer Ops Console.

Evidence
"THANK YOU for putting this in Web first! We have been using Resolve AI, but… we can forget about Resolve now and just stick with you." — Enterprise media and telecom customer

The agent running an investigation in the Ops Console side panel — the surface research pointed to, and the one I designed for reuse across the web product.
Every investigation runs the same triage path, and what it learns feeds back in — so the agent gets better on your systems, not generically.
Decision 02

Customer context matters more than prompts or general knowledge

Two of the foundational ideas here were mine, and they were table-stakes for the agent to be successful.

  1. Incident- and service-specific memory files the agent can add to over time, so it gets better on your systems rather than generically.
  2. Getting the customer's service runbooks up front, so it starts an investigation with the same context a new hire would get.

Both turned a generic reasoning engine into something that knew where it was.

Evidence
"It provides step-by-step procedures for isolating issues, similar to what network engineers actually do in practice." — Managed IT services customer

Decision 03

End every message with the context to make the next move

An agent that narrates without directing leaves the responder to translate prose into action under pressure. I made highly directive next steps a requirement of every message the agent sends — so at any point you can act, check its work, or take over, without reading back through the transcript.

Evidence
Conversation analysis showed where people dropped or stalled: the moments the agent explained without telling them what to do next.

Every message closes with next steps the responder can act on, one-click actions, and an offer to ingest a runbook and remember it for later.

04 — Outcome

It launched as an experiment, one of five agents PagerDuty shipped around the same time. We had the right approach to the right problem, and our customers responded: it became the most engaged-with agent at PagerDuty. The company reorganized around it, marshaling new resources onto a project that started as a bet and became PagerDuty's top-line company OKR, 1 of 3, with four delivery teams' roadmaps rewritten to support it.

"Using AI for the right tasks to solve the right problems is actually the real legit meaning of how you should use AI. And I think PagerDuty's got it right."

— Ghandi Kumar, Principal Incident Commander, Platform Engineering, Twill

"We've reduced the number of people needed to resolve major incidents by 25% in just two months."

Andy White, Director of Technology Operations, Checkout.com

05 — SETTING THE STANDARDS

Now every agent launch starts here.

"You're... paving the way for agents in the web product. Setting the standard for what a 'good' agent launch looks like across Slack and Web."

Drew McKinney, Design Leadership, PagerDuty

After GA the agent had four delivery teams behind it, each with its own PM and its own goals, and seven teams sat in my pillar, many with projects I wasn't designing for directly. I couldn't be in every room, so anything I needed to hold across all those teams had to live as a pattern or a process rather than a decision I blocked.

SHARED PATTERNS Agentic and conversational patterns for trust, honest expectations, and predictable outcomes
Components I designed the agent side panel for reuse from the start; it's now the basis of PagerDuty's cross-platform agent experience
Rituals Design crits, designer working sessions, and cross-functional share-outs. I started them after duplicate work surfaced in a sprint because nobody fully owned the seams between teams' domains
Process AIOps was the first department on the pillar-lead model, and the cadence we set is being replicated across PagerDuty's design org
People I mentored two senior designers and an intern, and supported my designers through promotion

"She took the time to walk me through the context, explain the problem space, and make sure I was set up to contribute effectively… she elevates the people around her, which is a quality I strongly associate with leadership."

Anastasiya, Product Designer, PagerDuty
Next: Event Orchestration →
All work → Contact
Lehua Gray · Staff Product Designer
Decision 01

The agent running an investigation in the Ops Console side panel — the surface research pointed to, and the one I designed for reuse across the web product.

Decision 02
A monitor breach triggers the SRE Agent, which checks status, pulls traces, summarizes patterns, and checks code changes to produce a diagnosis and next steps; a green arrow shows those learnings feeding back into the agent.

Every investigation runs the same triage path, and what it learns feeds back in — so the agent gets better on your systems, not generically.

Decision 03
An SRE Agent analysis in Slack, ending with a Next Steps list and one-click actions to upload a runbook, check related incidents, or get change events.

Every message closes with next steps the responder can act on, one-click actions, and an offer to ingest a runbook and remember it for later.