Mario Saved the EU but Broke My System

Hamed Silatani
|
Tags:
Blog
Incident Management
Outages
Simple thumbnail - conference table & headline reading 'Eurozone crisis live: Mario Draghi vows to save the eurozone'
IN THIS ARTICLE

Ready to make incident response your competitive advantage?

See how Uptime Labs builds provable, scalable incident response capability across your organisation.

Fourteen years ago (almost to the day), I experienced an incident that I unlocked new insight into when I read  Gary Klein's Sources of Power. The following story illustrates why experience, not training, is what develops incident response expertise.

July 26, 2012: The Setup

I was working at a trading firm. Everyone in the office knew the Eurozone was in trouble and that an announcement was coming. We'd prepared meticulously: load-testing all our trading systems to 10x their normal capacity, running through every scenario we could imagine.

Around 11:30 BST, Mario Draghi made his announcement at a global investment conference in London: the ECB would do ‘whatever it takes’ to save the Eurozone. It was a moment of relief for investors. The market started to move rapidly, and we could feel the pressure building on our systems.

Then, 30 minutes later, everything broke.

The First Halt

Around noon, everyone on our vast open-plan engineering floor suddenly jumped up from their desks. The horror on people's faces was unmistakable. Even the Chief Operating Officer rushed up to the engineering floor on the 6th level, which, in the normal order of things, is never a good sign.

The system had simply stopped & nothing was working. More specifically, load balancers connections had shot up, web servers had exhausted all their connection and middleware services stopped processing. There were no warning signs; no alerts from our expensive monitoring tools. It was utterly inexplicable. A couple of minutes later, it resolved on its own, and everything came back to normal.

For me, there was a sigh of relief. But the senior engineers who'd been in the industry long enough to know better looked visibly shaken. They knew something: the worst type of incidents are the ones that resolve on their own. That means you don't know what you fixed, and it will happen again.

The Pattern Emerges

Sure enough, at 12:30, the exact same outage happened. This time, people came running back from lunch, panic rising. When it resolved again a few minutes later, nobody felt easy. We all knew: judging by the pattern, the next hit would come around 1pm.

And that was the critical moment. One hour before the US market open at 2pm: the busiest, most consequential time of the trading day when American traders would come online and react to Draghi's announcement. So if the system failed then, the consequences would be catastrophic.

The Geeky Joke That Saved Everything

While everyone was frantically scratching their heads, trying to explain what was happening, one engineer made an offhand comment. He said it looked like a ‘major garbage collection’ - a Java thing related to memory management. A stop-the-world garbage collection event where the entire JVM pauses.

He laughed. No one else did. But in that moment, something shifted. In the absence of any other lead, everyone in the room latched onto this intuition. It was a universal moment of singularity: we had something to pursue.

Here's what I didn't understand then, but understood after reading Klein: that engineer didn't randomly guess. He saw a pattern. His experience with Java systems allowed him to recognise something that novices would have completely missed. This is the core of what experts see that the rest of us miss during incidents, in the words of Klein:

‘Intuition is when we use our experience, and the patterns we have learned, to rapidly size up situations and know how to respond without going through deliberate analysis.”

Following the Thread

We had over 1,000 services running across the platform. All the critical ones had garbage collection monitoring and alerting in place, and nothing was alerting. So we shifted focus to tier 2 and tier 3 services: the non-critical systems that might not have been monitored rigorously or might not even have GC logging enabled. We wouldn't have any way of knowing if something happened to them.

At 1pm, the third outage hit. Now the war room was fully formed. The COO, the Head of Trading, compliance officers - people who rarely set foot on the engineering floor were all there, huddled together, discussing how to respond to clients, how to manage the incoming calls, how to prepare for what would happen at 2pm when every major client would be online watching their trades.

Ten minutes later, an engineer buried in the tier 3 logs noticed something: a tier 3 service had been performing major garbage collections on a suspicious schedule. It was such a low-importance service that normally no one would have even looked at it. The only reason it got flagged was the time correlation with our outages. But critically, no one could explain how a garbage collection in that service could possibly affect the entire trading flow. We had no proof of causation. But we had no other leads, and we were out of time.

Decision Under Extreme Pressure

We quickly huddled and worked out our options. What could we do?

  • Add more memory?
  • Restart the server (rolling or full)?
  • Shut down the service completely?
  • Cut it from the load balancer?

What fascinated me, and what I only understood years later reading Klein, was how the senior engineers evaluated these options. In less than a minute, they ran through each one mentally, simulating what would happen if we took that action i.e. "If we add more memory, the next garbage collection will just be longer. If we shut down the service, the messaging broker it consumes from will pile up with messages and get flooded. If we isolate it from the load balancer..."

They settled on isolation. Cut it from the load balancer. It was reversible, surgical, and - as we later confirmed after the crisis - it was the only option that wouldn't have made things worse.

The Final Wait

By the time we organised ourselves to isolate the service, we hit another episode at 1:30pm. The system halted again. A few minutes later it recovered, but we were running out of time. The business was already preparing contingency plans: how to apologise to the market, how to think about compensation, how to inform clients if this happened at 2pm.

Then, about fifteen minutes before 2pm, we finally managed to cut the service from the load balancer.

We had 15 agonising minutes to wait and hope for the best. We couldn't do anything else. It felt impossible.

Then 2pm arrived & nothing happened. The sense of relief across the floor was overwhelming.

Why This Story Matters 14 Years Later

I didn't fully understand why this story stuck with me until I read Gary Klein's Sources of Power. Suddenly, everything crystallised with new insight:

Klein identifies 4 cognitive powers that emerge under the kind of pressure we experienced: extreme uncertainty, no clear clues and time pressure. These are the sources of power that separate experts from novices.

Intuition: That engineer's joke about garbage collection wasn't a wild guess or a moment of whimsy. It was pattern recognition: the ability to rapidly size up a chaotic situation with minimal information. Experts see patterns that novices don't even know to look for.

Simulation: In less than a minute, senior engineers mentally ran through each option one at a time, simulating outcomes. ‘If we do X, what else might happen downstream?’ This kind of scenario thinking was critical because we couldn't test anything or wait and see. We only had minutes. Novices would need to deliberate; experts could just simulate.

Metaphor: Previous garbage collection incidents informed their reasoning. They remembered metaphors and examples where adding memory made things worse, which helped them discount those options quickly. Experience creates a library of patterns.

Storytelling: This story stayed with me for 14 years. It's the kind of vivid incident that surfaces in memory under future pressure, making knowledge available not just to me but to anyone who hears it. This is how expertise spreads, and this is precisely why incident response training must evolve beyond classroom exercises.

John Allspaw, Founder and Principal, Adaptive Capacity Labs 10mo ·  There are only two ways people learn from incidents:  Personal, first-hand experience Via the experience of others (i.e., vicarious learning)  We cannot influence how or when #1 happens. We CAN influence how and when #2 happens.  Vicarious learning is the only way effective learning from incidents can scale beyond one person.  Creating conditions where this happens can be difficult. However, it's possible as long as there's broad recognition in the organization that...  a. Effective post-incident analysis means building the richest understanding of the event for the broadest possible audience.  b. The quality of post-incident analysis needed for (a) requires skill and expertise, in the same way experienced software engineers can produce higher-quality code more efficiently compared to when they first started.  c. Most organizations do not have this expertise, but these skills can be learned and improved. (This is what we do.)  Incident are being prevented all the time...in many cases, over 99% of the time! This takes effort, skill, and expertise.  Enabling the broadest audience to learn something they didn't know before also takes effort, skill, and expertise.

John Allspaw’s LinkedIn post advocates for systematic storytelling for vicarious learning - specifically, in the form of post-incident analysis.

The Core Principle: Experience → Expertise

Here's what I can articulate now, with Klein's help: the difference between an expert and a novice isn't innate brilliance. It's accumulated experience, whether real or simulated, that builds these 4 powers.

Most incident responders can't wait decades for real incidents to teach them. That's where simulation comes in. When you give a team carefully designed simulated incident experiences, you're not just teaching them facts. You're building their intuition, their ability to mentally simulate options, their metaphorical library of past incidents and their capacity to tell stories that will resurface under pressure.

The things you gather from each incident, such as the visceral details, the decisions made, the outcomes, these stay with you forever. They become the patterns your brain recognises instantly. That's where the power is. That's how we reduce recovery time: not through static learning, but through building expertise one incident at a time.

Hamed Silatani

Hamed is the co-founder and CEO of Uptime Labs. He has 20 years of experience in engineering leadership, reliability engineering and IT operations. Having spent the majority of his career at the sharp end of incident response in financial services, he's looking to help all companies master the unexpected.

Share this post

Ready to make incident response your competitive advantage?

— Chris Voss

See how Uptime Labs builds provable, scalable incident response capability across your financial services organisation.