Incident Response Is Culture, Not Procedure
Incident Response Is Culture, Not Procedure
When a system breaks, most organizations run to the same place: procedures, runbooks, and checklists.
But in real time, a simple thing becomes clear: what determines the outcome isn’t the document - it’s the people operating it.
The Illusion: “If We Write a Good Procedure - We’ll Be Ready”
Incident response is sometimes seen as a process problem. The assumption is that if we define who does what and in what order during a failure - the system will behave as expected.
But real incidents don’t arrive according to a script.
In a real Incident:
- Information is partial
- Time is pressing
- And the impact is unclear
In such a situation, a procedure is at best a starting point - not a rescue mechanism.
What Actually Gets Exposed During a Failure
An Incident isn’t just a technical failure. It’s a moment when the organization gets exposed.
During a failure, it becomes clear:
- Who feels responsible even when it’s “not theirs”
- Who’s afraid to make a mistake and so stays silent
- Who’s willing to act without full certainty
These aren’t traits of a system. They’re traits of a culture.
Blame or Learning - the Quiet Choice
Most organizations don’t say “we blame.” But they ask questions that feel that way:
- Who pushed the code?
- Why didn’t you test it?
- How did this pass review?
These aren’t learning questions. They’re attack questions.
And the result is almost always the same:
- Information doesn’t get shared
- Risks get hidden
A Simple Example: The Same Incident, Two Organizations
In one organization, an engineer notices a strange sign. They’re not sure it matters, and hesitate to report it. Time passes, and the damage grows.
After the event, people ask: “why didn’t you say something sooner?”
In another organization, an engineer notices a similar sign. They report it immediately - even without certainty. The event gets handled early, and the damage is small.
The difference wasn’t technology. It was the sense of safety to speak up.
Incident Response as an Organizational Capability
A responsible organization doesn’t just ask “what procedure should we write.”
It asks:
- Is it okay to make a mistake out loud
- Can we stop damage before we understand everything
- And does learning matter more than justification
Procedures can support this - but they don’t replace culture.
The Bottom Line
You can teach people procedures. It’s much harder to build a culture.
But only the right culture lets a system:
- Survive failures
- Improve from them
- And not repeat them again and again
Good Incident response is built long before the Incident.
Looking Ahead
After seeing how organizations act under pressure, we arrive at the most mature question:
when to stop. When to give something up. And when “good enough” is the right decision.
In the next post: Good Enough Engineering - and why perfection is sometimes the biggest risk.
📚 More in this Series: When the System Is Already Running
- Part 1 Production Is the System's Point of Truth
- Part 2 Latency as an Organizational Problem, Not a Technical One
- Part 3 When Metrics Lie
- Part 4 Deploy Is a Dangerous Event
- Part 5 Gradual Rollout: Why "Gradually" Isn't Always Safe
- Part 6 Backward Compatibility as a Long-Term Commitment
- Part 7 When a System Reflects Organizational Structure
- Part 9 Good Enough Engineering
- Part 10 Over-Engineering and Under-Engineering: Two Sides of the Same Mistake
- Part 11 Mature Engineers Don't Seek Control