Field note
The Incident Isn't Over When the Fire's Out
There's a specific kind of relief when systems come back online after an incident. Access is revoked, the box is isolated, the dashboards are green again. It's tempting to call that the end. It's actually the halfway point.
The gap that costs you later
Containment answers "how do we make this stop." It rarely answers "how did this start," and those are different questions with different amounts of urgency attached. Once the fire's out, the pressure to move on is enormous — and that's exactly when the evidence starts degrading. Logs roll over. Memories of "wait, I noticed something weird three days ago" get fuzzier. The window to actually reconstruct the timeline is much shorter than the window to contain the incident.
What closing the loop actually looks like
- Reconstruct the timeline while the logs still exist, not after retention policy has quietly deleted the evidence.
- Write down what you've ruled out, the same discipline that saves you time debugging any hard problem — so nobody re-investigates the same dead end a week later.
- Run it blameless, but run it specific. "Human error" is not a root cause. The actual root cause is almost always a boring, chained misconfiguration that existed long before the incident — the same kind of thing that makes offensive engagements work in the first place.
- Verify the fix, not just the symptom. Confirming the box is clean isn't the same as confirming the access path that got them there is closed.
Declaring victory when the alert clears feels like progress. Real progress is closing the gap that let it happen at all — because that gap doesn't patch itself just because nobody's looking at it anymore.