Ready to make incident response your competitive advantage?
See how Uptime Labs builds provable, scalable incident response capability across your organisation.
Your engineering organisation is growing fast. Is your incident response growing with it?
Early-stage engineering is exciting partly because so much happens organically.
The team is small. People know each other. They know the codebase well, they know who owns what, and when something breaks, people naturally fall into their roles.
A few experienced engineers tend to know exactly what to do. Someone understands the infrastructure. Someone knows the affected service inside out. Someone naturally starts coordinating, and people just pick things up.
At this stage, you might not have an Incident Commander (let alone multiple). You probably do not need a detailed runbook. There may not even be much of a defined incident process, let alone a dedicated incident management or SRE function.
And that can be completely fine!
In fact, at this stage you often have a resilience superpower that you may not even realise you have.
There are fewer processes, less layers of coordination and fewer people involved in making decisions. The structure is relatively flat, information travels quickly and people can act without waiting for multiple approvals.
When something unexpected happens, the organisation can adapt remarkably quickly.
The challenge is not to replace that flexibility with process as you grow. It is to preserve it while making incident response work across a much larger group of people.
The whole response effectively lives in people’s heads. It is implied. People know the unwritten rules, they know who to bring in and they know how things normally get resolved.
The problem starts when you scale.
Growth turns experience into dependency
More engineers join. Alongside, the codebase expands. More services appear, systems become more intricate and more people enter the on-call rotation.
But quite often, incident response still relies on exactly the same people who made the informal model work in the first place.
That is where dependencies start to form.
The same experienced engineers get pulled into every serious incident. The knowledge of how to run incidents remains concentrated in one pocket of the organisation.
Meanwhile, newer responders might know their own service very well but have very little idea how to actually manage an incident.
They may jump straight into remediation without first assessing what is happening or how wide the impact is. Nobody clearly takes ownership. Customer communication gets forgotten. Internal stakeholders have no idea what is going on. People get pulled into the response without being brought to the same level of understanding.
Sometimes incidents are not even properly declared or shared internally in the first place. And you are still relying on the same experienced people to come in and sort it out.
Except those people now have a much harder job too. The organisation is larger, the codebase is bigger, services are more interconnected and they no longer understand every corner of the system in the way they might have done earlier.
There is another problem too.
As the organisation grows, those same experienced engineers tend to accumulate more strategically important responsibilities. They are leading teams, designing systems, making architectural decisions, mentoring engineers and working on increasingly important problems.
So you end up with a strange paradox: their time becomes more valuable and more constrained, while the organisation becomes even more dependent on them when incidents happen.
They have less time to spend on incidents, but they are needed in more of them.
You end up with two dependencies: you depend on the same people for their technical knowledge, and you depend on them because they are the only people who really know how to run the incident.
That is not scalable. It is not sustainable. And it is definitely not resilient.
It can also burn those people out.
If every difficult incident eventually lands on the same few shoulders, they are never really removed from the response. It could be 2pm, it could be 4am, but when things go seriously wrong, everyone knows who is going to get called.
At some point, that model stops working. Think of Al Sene’s simple-and-effective quote: ‘Scaling reliability requires a focus on technology, people, and process.’
Incident response has to become an organisational capability rather than something a few experienced people happen to be very good at.
Because what you ultimately want is a resilient organisation, not a robust one where chaos naturally ensues whenever a novel situation arises.
So how should you organise incident response?
There are a few different ways companies approach this.
Some build a strong central SRE, reliability or incident management function that takes significant responsibility for coordinating major incidents.
Others push ownership directly into engineering teams. If you own the service, you also own what happens when it fails.
And increasingly, there are hybrid models. Engineering teams own their services and the actual response, while a central SRE or incident management function owns the framework around them: standards, tooling, escalation paths, coaching, training and support.
There is not one universally correct model.
The right answer depends on the organisation, its systems, its risk profile and how engineering is structured.
But whichever model you choose, you eventually run into the same reality.
Incidents are sociotechnical events
As Vanessa Huerta Granda recently established, incidents are sociotechnical events.
In other words, we have become extremely good at building technology around incidents.
Alerting tells us something has gone wrong. Observability helps us understand what is happening. Incident management tooling helps coordinate workflows and bring information together. Automation can gather context, run diagnostic steps and increasingly perform parts of remediation.
AI will automate more of the investigation layer as well. And that should happen. We should automate as much of the technical side as we reasonably can.
But an incident is not just a technical problem. It is a socio-technical event. To quote Vanessa, ‘Incidents will often start as technical failures. But incidents don't usually stay purely technical.’
That’s because someone needs to:
- assess what is actually happening
- take ownership (recall Lorin Hochstein’s iconic quote about AGI)
- decide who to bring in
- bring the new people up to speed without derailing the people working on the problem
All of this will require trade-offs need to be made with incomplete information, all under pressure.
At the same time, customers, leadership and other stakeholders need to understand what is happening. Someone needs to manage those communications without constantly distracting the responders who are trying to restore the service and run the incident.
These are not secondary, soft or ‘nice-to-have’ skills. They are a core part of incident response, and they will show up every single time.
A playbook is not the same thing as capability
Processes and playbooks matter: they create consistency. They help define roles, escalation paths and expectations. They give teams a shared understanding of what good incident response should look like.
But no two incidents are alike, due to a number of variables:
- Which systems fail.
- Which people are involved.
- Which customers are affected
- Which order the information arrives.
- What time of the day.
- The level of pressure.
At some point, responders have to exercise judgement.
And when a service is down, customers are affected, the company may be losing money and everyone is watching the response, nobody wants to sit there mechanically working through a runbook line by line.
So if you’re the incident commander, you want the team to collaborate. You want them to swarm around the problem. Furthermore, you want people to know how to assess the situation, coordinate the response, make decisions and communicate clearly while everything is moving around them.
The technology will change. The tooling will improve. AI will change how incidents are investigated, and probably how software itself is built.
But incidents themselves will continue to be different every time.
The human skills needed to manage them well are remarkably transferable:
assessment! ownership! coordination! decision-making! communication!
The underlying failure changes.
Those skills keep showing up.
Real incidents are an expensive place to learn
Of course, you can learn these skills by responding to real incidents.
A lot of experienced responders have done exactly that. But it is an incredibly expensive way to learn.
Firstly, it’s expensive in terms of the business impact itself: downtime, customer disruption, lost productivity, reputational damage and lost revenue.
Then it costly from the human angle: learning how to coordinate a response, make decisions under pressure and communicate clearly for the first time while an actual production system is failing is hugely stressful. Overall, it’s a lot to put on someone. And it is not a sustainable way to build capability across a growing engineering organisation.
So the challenge is not simply deciding whether incident response should sit with an SRE team, a central incident management function or the engineering teams themselves. It is making sure that whoever is expected to respond is ready (and, even better, feeling confident).
Because resilience is not about having the perfect process. It is about having enough capability across the organisation that when something serious happens, you are not dependent on the same few experienced people coming in to save the day. To quote Hamed’s post about incident hero culture - ‘A resilient organisation does not create one legendary responder. It creates twenty calm, capable people who can step forward when needed’.
You have a resilient bench of responders and incident commanders who can assess what is happening, coordinate effectively, make decisions under pressure and communicate clearly while the service is being restored.
AI may change the technical side of incident response significantly. It may automate more investigation, diagnosis and remediation. But it does not remove the human side. Humans still have to manage the incident.
And the skills that allow them to do that well - assessment, coordination, decision-making, ownership and communication - are not going away as AI surges forward in the incident response space. If anything, they become even more important.
That is how you build a resilient organisation: one that can protect its people, its customers, its reputation and its revenue when things go wrong.


