Incidents are Leadership Moments

Joe Mckevitt
|
October 1, 2026
Tags:
Blog
Outages
Best Practices
IN THIS ARTICLE

Ready to make incident response your competitive advantage?

See how Uptime Labs builds provable, scalable incident response capability across your organisation.

Every outage ends eventually:

Slack is silent. The dashboards goes back to green. I book a retro into a calendar event. But long after the incident itself is fixed, one thing stays with the team: how you behaved while it was still live.

That is the fifth principle in my CTO series, and it might be the one that matters most. Incidents are leadership moments. They are not distractions from the roadmap, or disruptions to product velocity. They are not merely outages to be patched. They are the moments where your stated values either hold up or fall apart, in front of the exact team you are trying to build that culture with.

1. Incidents are a test you can’t fake

Ask any leader what kind of culture they want to run, and you’ll likely get the same answers:

Calm. Curious. Blameless. Supportive.

Those words are temptingly easy to say in a strategy session. They get much harder to live out at 2am, with customers affected and no fix in sight.

That gap between the values in the Google Doc and the behaviour in the incident channel is exactly why this moment matters so much. You cannot fake it under pressure. Whatever shows up when things break is the real culture, not the one you’ve written down.

2. Support over blame, even when it is expensive

Every company I have worked with, including mine, has inevitably had its share of incidents.

This is a vertical timeline of an incident. Circular markers sit along a dashed line, with each event's label placed alternately to the left and right. The colour of each marker shows severity. Yellow marks the deployment events, orange marks early degradation, red marks the active incident, and green marks the resolution.  July 17:  16:10 (yellow): Version 0.45.8 of Session was deployed into dev. 19:00 (yellow): Version 0.45.8 of Session, which contained the defect, was deployed to production.  July 18:  07:30 (light orange): The broker began to degrade because of out-of-memory (OOM) issues. 10:05 (orange): The first user report was received. 10:18 (red): The incident was formally declared as Sev-1. 11:09 (red): The root cause was identified as broker overload. 13:29 (red): A decision was made to recreate the broker for faster recovery. 13:49 (green): Service was restored, and full platform functionality was confirmed shortly after.

A historic artefact: a timeline of an incident Uptime Labs had last summer

Leaders use incidents as a flywheel for improvement. Culturally, they set the bar to learn from the incident. They take the lessons seriously, apply them across the org and give everyone touched by the incident real room to learn and improve.

That is easy to say when nothing is at stake. It gets much harder when the incident has a price tag attached: lost revenue, a damaged relationship, an uncomfortable call with a customer, etc.. In that moment, blame is the cheap option, but unrewarding in the long term.

Here at Uptime Labs, when an incident impacts our business, we use the post-incident to identify areas for improvement and make them concrete, actionable, and focused on follow-through. Consciously choosing a constructive, rather than an accusatory approach, is what makes a blameless culture real instead of performative (my colleague Karan writes very well about this here).

3. Your customers are grading you too

Incidents are inevitable: it’s how you react to them that's key. Customers appreciate it when you hold your hand up and take accountability. What they will not forgive is silence, spin or a company that treats them like they should not have noticed.

How you communicate after an incident is a direct test of the relationship. Get it right, tell people what happened, what you are doing about it, and when they can expect an update, and you can come out of an outage with more trust than you had going in. Get it wrong, and the technical fix stops mattering. The story people remember is the excuses and cover-up, not the outage.

So it’s an opportunity for you as an organisation to show your maturity and to build trust with your customer, because how you react is very important and how you communicate that to them directly. So potentially over the long term, it actually helps you to build up more trust with your customer. They can see you actually becoming more resilient and taking the lessons learned from the incident.

A collage of news headlines cut from different publications, each in its own typeface, arranged around a central title. The headlines all describe companies apologising for, or compensating users after, service outages. The title sits in the middle in large, bold, dark blue sans-serif text on a pale background: "Incidents Will Happen. Are You (Actually) Prepared?" The word "Will" is wrapped in asterisks for emphasis.  Headlines above the title:  "'Deeply sorry,' CrowdStrike boss apologises for global IT outage", in white text on a black banner dated 24 September 2024. "Sony promise 5 days extra PS Plus as compensation for PSN outage." "Cloudflare CTO apologizes to the internet as a whole after global outage." "Facebook Apologizes for Mass Outage, Insists 'No Malicious Activity' to Blame", with the subheading "The service disruption cost the social media giant an estimated $100 million in lost revenue." "Zoom apologises after partial global outage."  Headlines below the title:  "Telstra CEO says she is 'deeply sorry' for nationwide outage." "Fortnite Dev Announces Free Rewards to Apologize For Downtime." "Google issues apology, incident report for hourslong cloud outage." "BlackBerry Outage: RIM Apologizes, Says Service Returning", with the subheading "Says service restored, but massive backlog of messages." "Roblox CEO apologies after three-day blackout." "AWS apologises for 14-hour outage and sets out causes of US datacentre region downtime.

4. What your team actually learns from watching you

Staff do not learn your values from the company handbook. They learn them by watching what you do the moment something goes wrong. Every incident is a live demonstration, whether you plan it that way or not.

When I walk into an incident, I try to treat it as the moment to prove the culture I claim to want, not an unwelcome interruption to it. That is not a performance. It is the only way the rest of the organisation gets to see leadership in practice instead of in principle.

For example, we had a recent incident whereby I had to step in and help out by addressing the action plan. Again, I emphasise: at no moment was there any sense of blame. It was all about driving improvements forward.

One of our core principles as a company is that errors and mistakes belong to the system, not people. Look at the circumstances that allowed the mistake to happen, rather than blaming the person.

Leading through the ones you cannot avoid

Incidents are not going away. They are a feature of running complex systems, not a bug in your process. On top of this, these systems are now outside our control with third parties, meaning that difficult incidents are a when, not an if.

As a leader your role is to make sure that when incidents do occur you are the ambassador of a blameless culture. Ensure that the lessons learnt from the incident are fed back into the company so that they can be used to improve your overall incident management. To quote John Allspaw, ‘Never waste a good incident…. it’s an unplanned investment opportunity.’

Joe Mckevitt

Joe is the co-founder and CTO of Uptime Labs. A passionate technologist and developer, he has 17 years’ experience in building and scaling high-performing products and teams. Also a marathon runner, he’s wired for high performance. He loves creating cultures of constant innovation, and coaching people to develop their full potential.

Share this post
Clear blue sky with scattered white clouds.

Ready to make incident response your competitive advantage?

— Chris Voss

See how Uptime Labs builds provable, scalable incident response capability across your financial services organisation.