'My Boss Wants Me to Pick an AI SRE Tool': Q&A at Incident Fest (Adaptive Capacity Labs)

Sam Salter
|
August 11, 2026
Tags:
Blog
AI & Automation
Best Practices
Incident Management
IN THIS ARTICLE

Ready to make incident response your competitive advantage?

See how Uptime Labs builds provable, scalable incident response capability across your organisation.

During Incident Fest 2026, our friends over at Adaptive Capacity Labs, John Allspaw & Beth Adele Long, answered questions about the relationship between AI & humans in our virtual ‘AMA Marquee’. Here are a selection of the questions & answers.

Q: What are the safe and helpful use cases for AI in incident response based on where technology is today? Real-life examples would be very helpful.

Beth Adele Long:

In terms of safety, I’m a proponent of read-only access during incidents. I was going to say “unless it’s a relatively trivial / low-risk scenario,” but any write access that’s powerful enough to be useful is also likely to be dangerous. And incidents are already confusing enough without having to unwind a bizarre decision that was implemented at AI speed.

With that safety caveat in mind, in long-running incidents, I can certainly see AI being helpful in the same the way it’s already being used during routine work: as a thought partner to help responders figure out what’s going on and explain current behavior. (J. Paul talked about exactly this in his talk, actually, and his point is really important. When AI predictions are bad, they degrade performance much more drastically than its good predictions improve performance.)

I would love to see AI tools helping responders better with pattern-matching, but again, how that information is connected and then presented to responders very much matters. I’d also love to see AI supporting incident commanders by helping them make sense of the organization itself — who’s the right SME? Who do we page? Who has been working on the current incident long enough that they’re probably burned out, and I should send them on a humanity break? These are aspects we don’t think about enough but that really matter to effective incident response.

Q: My boss wants me to pick an AI SRE tool to introduce into our org. Where do I start?

Beth Adele Long:

Oh boy, this is a tough one. I would start by getting clear about any contrasts between purported aims and real aims. By which I mean: how much is this a pragmatic request based on specific expectations (“As a leader, I believe AI SRE will help us do X, as measured by Y”) and how much is this actually motivated by something along the lines of “the board / my VP / someone in power says we need to be using AI more, and I need you to make me look good.” The more the latter factor is in play, the less room you’ll have to negotiate based on the actual benefit of the tools. You may just have to pick something and let it play out.

In either case, I recommend looking at how much an AI SRE tool supports integration into everyday work. Are they getting lost in the leftover principle that Stu talked about, promising they’ll do work with no intervention? Or are they making life easier for your SREs? The latter claim is easy to test: do a pilot and see what your engineers say. Operational types are notoriously blunt and usually overloaded, so you’ll probably get a fast and honest take whether the tools are helpful or just annoying.

Finally: good luck. This is a tough time to be evaluating tools that are still very much in flux and figuring out how to provide genuine value.

Q: AI has saved a lot of time in the incident review process: sifting through loads of data, nicely constructing the timeline and extracting patterns. It saves a lot of time doing conversations and interviews. I wonder if other folks have seen such time savings.

John Allspaw:

The METR study in 2025 on developer productivity helped shine some light on how the perception of time savings/spent can be different than the actual amount of time savings/spent.

(Before testing, developers guessed AI would make them 24% faster. After using it, they believed they were 20% faster. Turns out it was actually 19% slower.)

I’m fascinated by this topic, so I have questions for you as well as others:

When it sifts through data, what data does it dismiss as unimportant?

Since all timelines are opinionated (because they’re constructed from raw data in ways that make sense to the author of said timeline), same question: how does the AI choose between events to include and events to dismiss?

Q: When using AI in incident response, people frequently say that you have to second guess whether the AI is saying something sensible or not. Isn't that the same with humans?

John Allspaw:

Evaluating what your colleague has said while you’re both responding to an incident is clearly something happens, yep. Whether or not you’re “second guessing” what they’re asserting depends entirely on your experience with the person in the past, what they’ve said earlier in the response, how they described how they arrived at what they’re saying, etc.

I’d guess that many people with experience responding to incidents can think of other ways that a human coworker’s contributions might be different than an AI agent’s contributions during an incident…?

Beth Adele Long:

To expand on John’s remark about “your experience with the person in the past”: with humans we have a deep intuitive sense of how to judge someone’s trustworthiness. The engineer who’s been at the company for 5 years and deeply knows this system gets weighted differently than a new hire; the person who’s “often wrong, never in doubt” gets more skepticism than the quiet person who only speaks up when they really know what’s going on. AI usually falls into the “often wrong, never in doubt” bucket! And my experience is that it’s a lot more expensive to evaluate AI’s wall of text in an incident than to guess the trustworthiness of a colleague’s terse assertion. Humans tend to be more efficient at building a shared context as a group (thanks, evolution). So yes, the second guessing is always happening at some level, but how that process unfolds is different in important ways.

To see the full range of Q&As, you can explore Incident Fest here.

Sam Salter

Sam is the official editor at Uptime Labs, working closely with a global community of engineers and practitioners to surface real-world insights into incident response. Sam aims to help turn hard-won operational experience into clear, practical perspectives aiming to support how organisations prepare for, respond to and learn from incidents.

Share this post

Ready to make incident response your competitive advantage?

— Chris Voss

See how Uptime Labs builds provable, scalable incident response capability across your financial services organisation.