You've (Just) Had an Incident. What Next?

Tags:
Blog
Incident Management
IN THIS ARTICLE

Ready to make incident response your competitive advantage?

See how Uptime Labs builds provable, scalable incident response capability across your organisation.

I've run enough major incidents to know that the first hour rarely goes the way people expect. Here's what I've learned about what actually happens inside a team when something breaks, how teams often reflect, and how we think about accountability without losing a blameless culture.

The first hour of an incident: what's really happening

Logistically, after detection, people try to understand what's actually happening, then work out whether how to move forward. Note that incidents do not progress through stages in a linear way. You may come back to assessing severity or revisit assessment based on new information that emerge as incident progresses.

Similarly, it's also complex from an emotional perspective. Individual engineers often wonder if they broke something. Some people see a symptom and immediately form a theory i.e. a gut instinct about the cause - and then start hunting for evidence to support that theory rather than staying open-minded. Meanwhile, managers want constant updates, and what actually happens in Slack is information overload.

But in my view, the real challenge in the first hour isn't technical; it's cognitive overload. If you divide an incident into stages, the first part of the time should be owned by the incident commander, and their job isn't to solve the problem.It's to reduce chaos: create the space for people to pick up tasks and start investigating, make decisions explicit, and keep communication flowing. Those first minutes - or even hours, in a bigger incident - are never about fixing the system. They're about making sense of a surprising situation, recruiting people who can help and creating sufficient psychological safety that people will voice their theories and be explicit about certainty of premises of the theory. That's what separates a good incident commander.

A word about pressure

I think back to a story from earlier in my career that says more about pressure than any framework could. I was working through a major incident next to a colleague who was technically very strong. Our manager - who was visibly under stress - came up behind us and demanded we check the logs. The pressure of being watched and barked at was so distracting that my colleague forgot the basic syntax of the view command. Our manager ended up spelling it out: "v–i–e–w, and then the file name."

It's a small moment, but a telling one. Technical skill isn't the bottleneck when the room is hostile. A junior engineer can have all the right preparation, the right runbooks, the right paired senior and still fold if the environment around them is built on pressure and intimidation rather than support.

What teams could overlook after

People naturally want to answer "why did this happen?" Uncertainty is uncomfortable, and immediately after an incident there's a lot of it. So people jump to ‘why’, but I think that's the wrong question, because it sends you straight towards root-cause analysis.

As we explored in The Technical Foundations of Incident Response, the right sequence is: can we stop the customer impact - stop the bleeding? Then, can we stabilise the system? Then, can we preserve the evidence? Only then do you investigate - that's where the actual learning happens.

However, when a team implements a change and encounters errors, fixation can quickly set in. They observe memory errors and become convinced it's a memory leak, focusing all investigation efforts on proving that diagnosis. The problem is that this fixation causes them to lose sight of the primary incident response objective: restoring service. During an incident, multiple decision paths are available (rolling back the change, restarting the service, modifying configuratio) each with different risk and learning profiles. By pursuing only one diagnostic avenue in parallel, teams sacrifice their ability to understand what actually happened. Gathering information before committing to a single response path, and considering what each approach might reveal, is as important as speed in bringing systems back online.

The second accident is making several big decisions in parallel: restarting, scaling up, changing configuration, with different people acting independently. I wouldn't say the process needs to be strictly sequential, but it makes it much easier to asses impact of each change if you try one thing, gather information, and only then move to the next. Doing three or four things in parallel might bring the system back faster, but it destroys your ability to learn afterwards what actually happened. Preserving the evidence is just as important as restoring the service.

Accountability without losing a blameless culture

A ‘no-blame’ culture in incident management should not mean removing human accountability or glossing over the decisions people make during crises. Rather, it means resisting the urge to simply label events as ‘human error’ and moving on. Every incident involves decisions made by individuals who bear responsibility for those choices, but understanding why they made them is critical.

The goal is to reconstruct the context in which decisions were made: what information was available at the time, what pressures and constraints existed, and what risks seemed apparent or hidden. When we skip this deeper analysis to avoid blame, we sacrifice the insights that could prevent future incidents. Accountability and learning are not opposites; they work together when we focus on understanding the decision-making environment rather than punishing the decision-maker.

A good postmortem embraces the human element; one of the the key skills of Postmortem (incident review ) is to conduct in a way that no one is uncomfortable. Part of this means avoiding reducing a complex failure down to one person's mistake. That's what a blameless culture actually means: not pinning a complex failure on one person or team, as we discuss in our regulatory incident response piece. Accountability, on the other hand, is about improving future outcomes, not assigning guilt.

So in a postmortem, the questions focus should be on, "What set of circumstances led to the incident? What information was available to the human operator at the point of the decision making? Why that decision made sense to them?" .

Definitely  not "who did it, or why did they do it?" If someone skipped a checklist and that caused a major issue, the question isn't "why did they skip it" - it's "was there pressure that made skipping it possible? What barriers should have stopped that?" It's about whether the system allowed the mistake, not about the person. Accountability, meanwhile, is about owning the future outcome - again, without assigning guilt.

Who owns the learning?

It's the incident commander's responsibility to make sure the learning happens, but I don't own all of the learning myself. The engineering team should own the technical timeline. Monitoring and operations should look at the alerts and the response coordination. Customer support should explain the customer experience during the incident. My job as incident commander is to make sure all of that comes together.

One of the most valuable questions I ask isn't "what failed?" It's "what made the incident harder to resolve than it needed to be?" That's the better question, and it usually surfaces poor documentation, confusing dashboards, unclear ownership, missing alerts and communication gaps. That's what comes out of these discussions.

Add Uptime Labs to your post-incident learning

None of this is instinctive. Reducing chaos before fixing the system, resisting the pull towards "why", asking what made an incident more difficult to resolve rather than who's to blame - these are habits people get better at by practising them, not by reading about them and hoping they’ll stay in your head while everything is on fire.

That's exactly why I helped build Uptime Labs. Our drills put your team through the same ambiguity, incomplete information, and communication pressure a real incident throws at you, minus the real-world stakes. You find out how your team actually behaves in that first hour before it costs you a customer. Each drill comes with a personalised report designed to support ongoing skills development.

If you want to see what that looks like in practice, try a drill or get in touch to book a demo.

Share this post

Ready to make incident response your competitive advantage?

— Chris Voss

See how Uptime Labs builds provable, scalable incident response capability across your financial services organisation.