So You Wrote The Post Mortem. Now What?

Most people treat a post mortem like a box to check. They write it, hit send on an email or post it in a shared drive, and move on to the next incident. That is why almost nothing ever changes after a failure happens. The document sits there gathering digital dust while the same class of problem shows up three months later in a slightly different form. The actual next step is less glamorous than people expect. It involves taking every action item from that report and putting it into a tracking system where someone has to answer for it. I have seen teams spend six hours agonizing over blameless language and then ship three bullet points with no owner assigned and no due date. That is not a process. That is a team building exercise that happens to produce a PDF nobody reads. Here is what I actually do after I close the writing phase. I dump every recommended action item into a dedicated project management board or ticketing system. Not a separate spreadsheet. Not a comment thread. Actual tickets with a single owner and a target completion date. If the incident revealed a gap in our alerting thresholds, that becomes a monitoring ticket. If a runbook was missing or stale, that becomes a documentation ticket. If something needs code work, it goes into the engineering backlog with a clear acceptance criterion.

I also schedule a retro or follow-up meeting at a fixed point in the future. Usually two weeks out. At that meeting we look at the tickets, not the report. We ask whether each one moved, stalled, or got deprioritized. This two-week cadence matters because it catches the momentum drop before it becomes permanent. I have seen action items die quietly when the follow-up meeting got pushed to next quarter by a busy sprint planning session. Once that happens, the whole exercise was wasted effort.

Tracking And Closure Is Where Most Teams Break

There is a specific edge case that always catches people off guard. A few years back I was dealing with an outage where the root cause traced back to a configuration drift between our staging and production environments. The post mortem recommended we implement infrastructure-as-code validation to catch those kinds of mismatches automatically. Good recommendation. Terrible execution path because nobody owned the actual implementation. I ended up creating a lightweight checklist for action item triage that I force through every post mortem before it leaves my desk. The checklist has four fields for every single item: the exact deliverable, the person responsible, the definition of done, and the date we will check progress. Without all four fields filled in, the ticket does not exist in our system. It sounds rigid but it prevents the classic ghost action item where everyone assumes someone else is handling it. In that configuration drift incident, the fourth field caught the problem early. Our first scheduled check-in revealed the ticket had been sitting unstarted for eleven days because three different engineers all thought a fourth person was picking it up. We reassigned it and it shipped within a week. This approach has a real downside though. It adds overhead. For smaller teams or less severe incidents, the formal tracking can feel like overkill. If you have a team of four people and the incident was a five-minute DNS misconfiguration, you probably do not need a dedicated ticketing workflow. A quick Slack summary with one or two action items and a verbal commitment to review them next standup is usually enough. The formal process scales with incident severity. Minor blips get lightweight treatment. Major outages get the full tracking pipeline.

Get the Full Details

Electronegativity and Bond Polarity - Chemistry Steps
Electronegativity and Bond Polarity - Chemistry Steps

Another counter-intuitive thing most teams miss is that some post mortem action items should deliberately not become tickets. If an action item is purely informational, like updating a diagram or sharing a lesson learned with another team, ticketing it creates false accountability pressure. These items should go into a lightweight read-only log or a simple shared document instead. Ticketing everything inflates your backlog and makes it harder to spot the items that actually need engineered work. I separate action items into three categories: engineering work that gets a ticket, operational changes that get a ticket, and informational items that get logged separately.

Measuring Whether The Follow-Up Actually Worked

You need a way to know if the post mortem process is improving. The metric that actually matters is the action item closure rate. Track how many recommended items from post mortems get completed within the target timeframe. If that number stays below sixty percent over three consecutive months, your process is broken and you are not fixing anything. The second useful metric is incident recurrence. Are the same class of failures showing up again in environments where you already wrote a post mortem about them? If yes, the action items from those earlier reports were never properly tracked or executed. There is also a timing consideration that affects everything. The post mortem and follow-up should happen while the incident is still fresh enough that people remember the details accurately, but delayed enough that emotions have cooled. I usually aim for forty-eight to seventy-two hours between the incident resolution and the report being finalized. Anything sooner and you are still guessing at some of the root causes. Anything later and people forget why certain decisions mattered. The workflow is not particularly exciting. Write the report. Extract action items. Put them in a tracking system with owners and dates. Schedule the follow-up. Review progress. Repeat. The gap between writing and tracking is where the value evaporates. Fill that gap consistently and you will notice a difference within a few months. Leave it open and you are just producing documents for archival purposes.