How to Train Incident Commanders

Peter Catack (Community Contributor)
|
July 21, 2026
IN THIS ARTICLE

Ready to make incident response your competitive advantage?

See how Uptime Labs builds provable, scalable incident response capability across your financial services organisation.

Incident commander training is the structured process of developing engineers to lead incident response. An effective programme develops three core competencies: communication, sociotechnical leadership, and cognitive load management. The goal is a bench of trained ICs who can coordinate a response under pressure, not a single heroic responder.

Incident command is one of the most consequential skills in an SRE organisation, and one of the least deliberately taught. Most teams default to promoting their most senior engineer into the IC role and hoping the skills transfer. In most cases, this is not enough.

This article covers what incident commander training actually involves: the specific competencies an IC needs, how to structure a training programme, the mistakes that stall IC development, and how to measure whether your training is working. If you already know what an IC does and want to build the capability across your team, you're in the right place.

For a complete framework you can implement directly, including IC selection criteria, programme structures, and readiness measurement, download the Training Incident Responders and Scaling Incident Response Teams guide.

Why Incident Commander Training Matters

Most organisations encounter the IC role when an incident forces someone into it. An engineer takes charge because nobody else did, or because they happen to be on call. The response works, or it doesn't, and the team moves on without examining whether the coordination held or fell apart.

The problem is that incidents rarely stay purely technical for long. A database overload or a bad deployment quickly becomes an organisational event: multiple teams investigate different parts of the system, engineers propose competing mitigations, leadership asks for updates, and support needs customer messaging.

Two systems are running simultaneously during every incident: the technical system and the human coordination system. The outcome depends on how both behave, which is why a technically simple issue can become severe when coordination collapses.

This is the gap that incident commander training addresses. The IC's primary job is not to solve the technical issue personally but to ensure the issue gets solved efficiently. They stabilise the response by coordinating people, information, decisions, and communication so that subject matter experts can focus on the technical resolution.

Technical depth helps an IC ask more valuable diagnostic questions, but the competencies that make someone effective in command are different: translating between audiences, noticing when the bottleneck is organisational rather than technical and maintaining a coherent picture of the incident under incomplete information. Those are learnable, and they require deliberate practice in conditions that approximate real pressure.

What Competencies Does an Incident Commander Need?

Before you can train an IC, you need to define what you're training for. The IC role sits at the intersection of three core competencies.

  1. Communication

Incident command is fundamentally a communication role. A dedicated IC frees subject matter experts to focus on troubleshooting the technical problem by handling the back-and-forth coordination across teams. The challenge is that ICs have to constantly switch audiences: engineers, executives, support teams, and customer-facing teams all need different things from the same information.

An effective IC can answer two questions clearly and quickly: what is happening technically, and why does it matter to the business? And they can frame those answers in a way that makes sense to each audience. That translation skill is what keeps stakeholders informed without pulling investigators out of their work.

  1. Sociotechnical leadership

A strong technical responder notices when the system has run out of capacity or hit an unexpected error. A strong IC notices when the bottleneck is not technical but organisational: missing ownership, wrong teams involved, unclear accountability, or communication gaps between groups who each hold part of the picture.

This is leadership in the context of the incident, not in the sense of org-chart authority. The IC is coordinating the organisational system that wraps around the technical one. They guide the response rather than dominating it. When an IC starts solving the problem themselves, they lose their view of the full incident, and the coordination suffers.

  1. Cognitive load management

Incidents generate fast, incomplete, and often conflicting information. An effective IC can track multiple investigation threads and maintain a coherent picture of where the response stands: what has been tried, what is being tried now, what the risks are, and what decisions have been made.

People who are strong in this area tend to be comfortable operating with ambiguity and partial information. When someone asks, "Where are we right now?" they can answer clearly. That ability to hold and reconstruct the narrative of the incident under pressure is one of the strongest signals of IC readiness.

These three competencies are what separate the IC role from technical expertise. An engineer can be exceptional at diagnosing system failures and still struggle in command because the skills are different. Training needs to develop all three deliberately, not assume they transfer from technical proficiency. What an Incident Commander is and what they're responsible for covers the role definition in more detail.

How to Build an Incident Commander Training Program

Define the target competency profile

Before any training begins, establish what "ready to IC" means for your organisation. This means setting proficiency targets against the three competencies: communication, sociotechnical leadership, and cognitive load management.

The target profile will differ by team size and incident complexity. A team running a single service with a small on-call rotation needs a different IC profile than a platform engineering organisation managing cross-team Sev-1s. Small teams combine roles; larger teams separate them. How incident management roles evolve as your team scales covers each growth stage.

Practise through simulation

Simulation is where IC competencies are built, not just described. Reading a runbook or attending a tabletop discussion tells an engineer what to do. Simulation creates the conditions for them to practise doing it under pressure: managing competing information streams, communicating across audiences, and making decisions when the picture is incomplete.

The gap that tabletop exercises leave open is pressure and measurement. A tabletop tells you whether engineers can describe the right actions. Simulation tells you whether they can execute them when the CEO is asking for updates and a second service starts degrading. That distinction matters because the competencies an IC needs (such as holding the narrative, staying calm, delegating instead of absorbing work) only develop through repetition under realistic conditions.

Game days serve a similar function. Teams inject controlled failures into staging or test environments and practise the full response cycle, including IC coordination. The value is in the surprise element and the tempo, though game days are more expensive to organise and harder to repeat at high frequency than platform-based simulation.

Shadow and reverse-shadow on real incidents

Simulation builds foundational competency. Shadowing transfers it to the complexity of real systems and real organisational dynamics.

Before any engineer carries the pager independently, they should shadow an experienced IC through real incidents. The goal is not passive observation. During a shadow rotation, the engineer follows along in all incident channels in real time, writes their own diagnosis hypothesis before the senior IC announces theirs, and debriefs after every incident.

A structured shadow rotation moves through three phases:

  • Observe: Attend real incidents as a silent, active note-taker. Focus on how the IC sequences decisions and manages communication, not the technical resolution.
  • Scribe: Take the documentation role. Maintaining the decision log forces active listening and builds the habit of timestamping every action.
  • Buddy IC: Take the IC role with a senior IC available to consult, but not to take over. This is the critical transition point where competency becomes confidence.

Reverse shadowing is the counterpart: the experienced IC observes while the trainee runs the response, intervening only if the situation requires it. This builds the trainee's ability to hold command without the safety net being invisible.

Coach after incidents

Every real incident a trainee IC participates in, whether as shadow, scribe, buddy, or solo, is a coaching opportunity. The debrief after the incident should cover how the IC managed coordination, where communication broke down, what decisions were made under uncertainty and how they landed. This is where the three competencies move from abstract to concrete: the trainee can point to a specific moment where they lost the thread of the incident, or where they successfully translated a technical situation for a stakeholder.

Coaching is what compounds the learning from simulation and shadowing. Without it, the trainee accumulates experience but not necessarily insight.

Measure readiness before assigning solo IC responsibility

The output of IC training is not a certificate. It is a measurable readiness signal that tells you whether an engineer is prepared to command a real incident without a safety net.

The readiness signal should be objective: competency scores from simulation, performance observations from shadow rotations, and a named sign-off from a senior IC. Subjective "I think they're ready" assessments are how teams end up with undertrained ICs in charge of Sev-1s.

Competency-level measurement matters here more than time-based metrics. An engineer who has completed ten simulation drills and three buddy IC rotations but still scores at Practitioner level on communication is not ready, regardless of how many weeks they've been in the programme. Readiness is demonstrated through behaviour, not elapsed time.

How to Identify IC Talent in Your Organisation

The people who will make strong ICs are often already inside your organisation. The challenge is recognising them, because the signals that predict IC effectiveness are not the same signals that predict technical seniority.

Our guide to Training Incident Responders and Scaling Incident Response Teams structures IC selection around the same three competencies the role demands. Rather than selecting by title, tenure, or prestige, look for behavioural signals in day-to-day engineering work that map to communication, sociotechnical leadership, and cognitive load management.

Competency Signals to look for You might hear them say
Communication Presents technical topics to non-technical teams. Writes clear incident summaries. Explains a complex bug in a standup without losing the audience. "The database is running out of connections, which means checkout will start failing for customers in the next few minutes."
Sociotechnical awareness Thinks about impact beyond the code. Loops in adjacent teams proactively. Considers how engineering decisions affect the broader organisation. "If this service goes down, the support team will get flooded." / "We should loop in the payments team before we deploy this."
Cognitive load management Thrives during chaotic debugging. Moves between logs, metrics, Slack threads, and competing hypotheses without losing track. Naturally reconstructs the incident narrative. "Here's where we are so far." / "These two symptoms may be connected." / "Let's recap what changed."

These signals are observable in simulation, which is one reason simulation-based assessment is more useful for IC identification than performance reviews or tenure. A 30-minute drill reveals coordination behaviour that a year of standups might not surface.

When hiring ICs externally, assess the same three competencies. Provide candidates with incident data and ask them to write updates for three different audiences: engineers, leadership, and customers. Run a tabletop or mock incident and observe how they delegate investigation, manage competing ideas, and communicate with stakeholders. 

For cognitive load assessment, show them an unfamiliar incident timeline and ask them to walk through what happened, what matters, and what should happen next. You're not evaluating whether they guess the right technical answer. You're evaluating how they run the response.

Common Mistakes in Incident Commander Training

Most organisations fall into predictable patterns when they first start building incident command capability. They come from treating the IC role the way the industry has historically treated it: as a procedural responsibility rather than a distinct skill.

  1. Making the strongest engineer the IC

This is the most common anti-pattern, and it creates exactly the problem it's trying to solve. Your strongest technical investigator is now writing updates, coordinating teams, handling stakeholder questions, and managing timelines instead of debugging the issue. You've removed your best diagnostician from the technical response and placed them in a role that requires a completely different set of competencies. Some engineers are strong at both, but the assumption that technical skill transfers to coordination skill is where programmes go wrong.

  1. Making whoever is on-call the IC

This happens when incident command is treated as an automatic duty attached to the on-call rotation rather than a trained role. Being on call does not mean someone has practised communication under pressure, prioritisation across competing workstreams, or decision facilitation with incomplete information. Expecting people to improvise those skills during a live outage creates coordination problems that compound the technical ones. This does not mean the on-call engineer can never be the IC, but for it to work they need training and the psychological safety to escalate when something is beyond their depth.

  1. Making the most senior person the IC

Seniority brings deep understanding of the sociotechnical system, which helps during incidents. But it also brings strong opinions and established authority that can distort the IC role. When the most senior person takes command, the dynamic shifts from facilitation to direction. Other responders defer rather than contribute, dissent gets suppressed, and the IC ends up driving the investigation rather than coordinating it. A strong IC guides the response without dominating it.

  1. Treating IC training as a one-time event

IC competency degrades without practice. The teams that stop drilling once the first cohort is through the programme are the same teams that find their coordination breaking down six months later when the situation is unfamiliar. IC training is a continuous programme: regular simulation, periodic shadowing refreshers, coaching after real incidents. Competency is maintained through repetition, not through a certificate earned once.

How Does Uptime Labs Support Incident Commander Training?

The training programme structure described in this article requires two things most organisations lack: realistic simulation infrastructure and a way to measure competency progression objectively. Uptime Labs provides both.

Engineers run progressive drills in a browser-based environment that mirrors production: real dashboards, real communication channel dynamics, and stakeholder pressure scenarios that test command behaviour, not just technical knowledge. The drills are designed to develop the three IC competencies specifically. Communication is tested through multi-audience update scenarios. Sociotechnical leadership is tested through coordination challenges where the bottleneck is organisational. Cognitive load management is tested through scenarios with competing information streams and incomplete data.

Every drill scores performance across 40+ behavioural metrics mapped to five competency categories: Identify Scope, Incident Mechanics, Internal Comms, External Comms, and Command Incident. Each category tracks proficiency from Practitioner through to Expert. After each session, performance is plotted on a radar chart so Heads of SRE can see exactly where each engineer sits relative to the team's defined readiness threshold, and where additional practice is needed before they carry solo IC responsibility.

That measurement layer is what closes the loop on IC training. Without it, you're relying on subjective judgment to decide when someone is ready. With it, readiness becomes a demonstrated competency level rather than an assumption.

For a complete framework on building a sustainable IC programme, including how to identify IC talent internally, common selection mistakes, and how to structure your IC team as your organisation scales, download the Training Incident Responders and Scaling Incident Response Teams guide.

FAQs: Incident Commander Training

Who should be an incident commander?

The IC role requires strong communication, coordination, and cognitive load management skills, not the deepest technical expertise. Many high-performing teams rotate the IC role across mid-level engineers who demonstrate these competencies. Selecting ICs by seniority often produces the opposite of what you want: an engineer too focused on the technical problem to stay in command mode.

How long does it take to train an incident commander?

There is no fixed timeline. Readiness depends on the individual's starting competency level and the frequency of practice opportunities. The Uptime Labs guide to Training Incident Responders and Scaling Incident Response Teams (written by Vanessa Huerta Granda) recommends a progression through shadowing, simulation, and buddy IC rotations, with solo responsibility assigned only when competency scores demonstrate readiness across all categories, not after a fixed number of weeks.

Does the IC need to be the most technical person on the team?

No. In fact, assigning your strongest engineer as Incident Commander removes your best investigator from the technical response. The IC's job is to coordinate the response, not to debug the system. Technical depth helps for asking sharp diagnostic questions, but the core of the role is communication, delegation, and maintaining situational awareness under pressure.

How do you measure whether incident commander training is working?

Measure competency, not just time metrics. Track proficiency scores across the specific behavioural categories the IC role demands: scope identification, incident mechanics, internal and external communication, and command behaviour. MTTR improvement is a downstream outcome of improved competency, not the measurement target itself.

Peter Catack (Community Contributor)
Share this post

Ready to make incident response your competitive advantage?

— Chris Voss

See how Uptime Labs builds provable, scalable incident response capability across your financial services organisation.