10 Things We Learned About Training and Scaling Incident Teams This Month

Sam Salter
|
September 25, 2026
Tags:
Best Practices
Blog
Incident Management
IN THIS ARTICLE

Ready to make incident response your competitive advantage?

See how Uptime Labs builds provable, scalable incident response capability across your organisation.

This month, we were lucky enough to host a workshop led by Vanessa Huerta Granda (Technology Manager for Resilience Engineering at Enova International, board member at Resilience in Software Foundation, formerly of Jeli.io) on training Incident Commanders and scaling incident response teams.

Here are the 10 takeaways we gleaned from Vanessa’s talk:

1. Incidents are sociotechnical events (not just technical failures!)

A database overload or bad deploy is only half the story. The moment multiple teams get pulled in, leadership starts asking for updates and Support needs customer messaging, you're managing a human coordination system running in parallel with the technical one. Overall, the outcome of an incident depends on how both systems behave.

"Incidents will often start as technical failures," Vanessa said in the workshop. "But incidents don't usually stay purely technical. Very quickly, they end up becoming organisational events... You're actually coordinating a group of humans who are interacting with a complex system under pressure."

2. The Incident Commander's job isn't to fix the problem directly. It's to make sure it gets fixed

Think orchestra conductor, not soloist. The IC doesn't play every instrument; they keep the whole performance coordinated so subject matter experts can focus on resolution instead of managing timelines and stakeholders.

"The IC role is not primarily about solving technical problems," Vanessa said. "It's about making sure that the problem itself gets solved efficiently."

Incident commander ≡ incident conductor

3. Strong ICs operate across three domains: People, System, and Business

  • People: managing stress, duplicated effort, and who's overloaded
  • System: synthesising what different teams see into one shared picture
  • Business: answering "how many customers are affected" and "how urgent is this" in plain terms for leadership
Article content

4. Mistake #1: Making your strongest engineer the IC

This pulls your best technical investigator off the actual investigation and hands them updates, coordination and stakeholder management. These are possibly skills nobody's actually confirmed whether they actually possess.

5. Mistake #2: Whoever's on-call becomes IC by default

Treating incident command as a procedural checkbox instead of a trained skill sets people up to improvise communication, prioritisation and decision facilitation live, under pressure (with no safety net).

6. Mistake #3: The most senior person takes over

Seniority helps with sociotechnical context, but it can also tip the IC role from facilitation into command authority. A good IC guides the response; they don't dominate it.

"The incident commander is helping facilitate decision-making," Vanessa said. "They are not the dictator."

7. Three competencies actually predict a good IC

Not title, tenure,or technical depth, but rather:

  • Communication across engineers, execs, support and customers
  • Sociotechnical leadership - noticing when the bottleneck is organisational, not technical
  • Cognitive load management - holding the mental map of a fast-moving, incomplete, conflicting information stream

In Vanessa's words, a good IC is "answering two questions at the same time: 'What is the system doing?' and 'Why should the business care?'" And on cognitive load specifically: "I like to think of this as maintaining a mental map of the incident."

8. Your best future ICs are probably already on your team

Watch for signals in everyday work: people who translate technical topics for non-technical audiences, who instinctively loop in the right teams before something breaks and who can calmly recap "here's where we are so far" mid-chaos.

"These people are often already in your organisations," Vanessa told us. "We just don't label them as ‘incident commander.’"

9. A sustainable IC programme needs three things: 1) structure, 2) the right people and 3) training

Structure can look like a dedicated IC team, domain-based ICs, or a volunteer rotation. You can choose based on company size and incident frequency. Then train through shadowing, reverse shadowing, game days and post-incident coaching.

The goal? Build the capability; don’t shoot for perfection.

"Regardless of the structure that you choose, the people acting as incident commanders need to know that this work is a priority," Vanessa said. "It's not something that they do on top of their real job. This is part of their real job."

Her preferred training progression:

  • Shadowing
  • Reverse shadowing
  • Running incidents with support
Article content

10. Success means incidents stop being heroic

When the programme works, response gets calmer, more coordinated, and more predictable. You're no longer relying on a few exhausted people saving the day every time. As Vanessa put it in her QCon talk: incidents are where engineers are made. The goal is always to build systems where no one has to be a hero.

Vanessa then closed the workshop with the line that really resonated:

"People with those skills exist. The challenge is recognising them and then giving them the support to succeed. Because the goal of incident command is not to create heroes - it's to create a system where no one actually has to be one."

Thanks again to Vanessa Huerta Granda for running this workshop with us earlier this month. If you want to go deeper, she covers all of this (and more on hiring and leadership buy-in) in her guide on training Incident Commanders.

Sam Salter

Sam is the official editor at Uptime Labs, working closely with a global community of engineers and practitioners to surface real-world insights into incident response. Sam aims to help turn hard-won operational experience into clear, practical perspectives aiming to support how organisations prepare for, respond to and learn from incidents.

Share this post
Clear blue sky with scattered white clouds.

Ready to make incident response your competitive advantage?

— Chris Voss

See how Uptime Labs builds provable, scalable incident response capability across your financial services organisation.