Getting Past the Surface With Thirteen Layers
Most people stop asking why after three or four iterations and call it a day. That leaves you with a fix for the symptom instead of the actual problem. The 13 Reasons Why Analysis pushes further because real root causes are rarely obvious. I learned this the hard way during a production line stoppage where our team kept circling back to "operator error" as the answer. By reason six we were going in circles, so we kept going. By reason ten we found a calibration drift in a sensor that had been misaligned during a maintenance changeover eighteen months earlier. The sensor was never re-verified after that service call. That is the whole point of doing this method properly. It is a structured root cause investigation technique that extends the classic 5 Whys framework by continuing the questioning chain up to thirteen layers deep. Each "why" builds directly on the answer to the previous one, creating a causal chain from the observable problem back through successive layers of contributing factors until you hit something fundamental. The output is a single continuous chain, not a branching tree or fishbone diagram. You write each answer as a standalone statement before moving to the next question. The method originated in lean manufacturing and systems engineering circles as a response to the premature termination problem inherent in shorter versions. Five iterations is often enough to reach a process gap, but insufficient for complex technical systems where organizational, material, and design factors layer on top of each other. Fourteen iterations is unnecessary in almost every real scenario, so thirteen became the practical ceiling used in most implementations I have encountered.
The Procedure
Start by writing the problem statement as a concrete, specific fact, not a vague complaint. "The pump seal failed at 0400 hours on Line 3" works. "Equipment problems are causing downtime" does not. The vaguer your starting point, the more your chain will drift into speculation by reason six or seven. Ask why this happened and write a factual answer based on evidence, not opinion. Then immediately ask why that answer occurred. Continue this pattern for thirteen iterations. At each step, pause and verify the answer before proceeding. If you cannot verify it with data, documentation, or direct observation, mark that link as uncertain and note what evidence would confirm or refute it. Here is a simplified version of what the chain looks like in practice:
- Reason 1: The valve did not close. (Verified by control system log)
- Reason 2: The actuator received no signal. (Verified by PLC diagnostics)
- Reason 3: The PLC output card failed. (Verified by spare card swap test)
- Reason 4: Capacitor C7 on the output card leaked. (Verified by visual inspection)
- Reason 5: The capacitor exceeded its rated voltage. (Verified by oscilloscope trace of startup transient)
- Reason 6: The soft starter was bypassed during commissioning. (Verified by commissioning notes)
- Reason 7: The soft starter manual was not followed. (Verified by work order review)
- Reason 8: The commissioning engineer did not have the manual. (Verified by HR records showing no training log entry)
- Reason 9: Training records were not updated after the SOP revision. (Verified by document control audit)
- Reason 10: The document control procedure has no enforcement mechanism. (Verified by policy review)
- Reason 11: No one was assigned ownership of procedural compliance. (Verified by org chart and job descriptions)
- Reason 12: The maintenance department never hired a dedicated documentation coordinator. (Verified by staffing history)
- Reason 13: The role was eliminated during the 2019 restructuring. (Verified by finance records)
At reason thirteen you are typically pointing at an organizational decision made years ago. That is the actual root cause worth fixing. Thirteen iterations sounds thorough, but the method has real limitations. The biggest one is that it assumes a single linear causal chain. Real failures usually have multiple contributing paths that converge on the same symptom. When two separate issues caused the same outcome, the 13 Reasons Why format forces you into choosing one path and ignoring the other. I ran into this on a wastewater treatment plant incident where both a software bug and a mechanical wear issue independently contributed to the failure. The chain became a mess because I was trying to thread two unrelated problems through one sequence. I had to restart the analysis twice to map both paths separately before I could see the full picture. Another limitation is verification at depth. By reason eight or nine, you are often asking about policy decisions, staffing choices, or budget allocations from years ago. Getting a reliable answer requires records access that many organizations simply do not grant to frontline engineers. I spent three weeks chasing a procurement memo that turned out to be in a different division's filing system. If your organization lacks document transparency, the chain will degrade into guesswork past reason six and you should stop rather than pretend precision exists where it does not.
Get the Full Details

The method also breaks down when the problem is inherently random or stochastic. Equipment fatigue lives on probability distributions. Using a deterministic why-chain for a failure mode governed by statistical variance gives you a false sense of certainty. In those cases, use Weibull analysis or fault tree analysis instead and reserve the 13 Reasons Why for situations where human process, policy, or design decisions created the conditions for failure.
Practical Workarounds I Use
When I hit an unbreakable organizational wall, I stop the chain at the last verifiable link and switch to a countermeasure mapping exercise. Instead of pushing deeper into unknowable territory, I document what we do know and work backward from the deepest confirmed reason to identify interventions. This usually takes thirty to forty-five minutes and produces an action list that is actually implementable. Trying to reach reason thirteen when reason ten is blocked by missing records just wastes time and produces an unreliable chain. I also run a parallel chain for the opposite direction. After building the standard forward chain, I ask "what would have prevented each answer?" This flips the analysis into solution space and often reveals that the most effective intervention sits at reason three or four, not reason thirteen. Organizations tend to fixate on the deepest root cause because it feels the most satisfying, but the cheapest and fastest fix is often much higher up the chain. A reworked work order template at reason eight might prevent this exact failure class across the entire site for less than two hundred dollars in engineering time.
When to Use It and When to Walk Away
Use 13 Reasons Why Analysis when you have a discrete, documented failure event with an accessible team, available records, and sufficient time. Two to four hours is realistic for a properly done chain on a moderate complexity system. Anything more involved than that and you should bring in a formal RCA process with a trained facilitator. Walk away and use a different method when the failure is ongoing and recurrent without a clear starting event, when records are systematically destroyed or unavailable, when the problem involves complex adaptive systems with no single causal path, or when the organization treats root cause analysis as a compliance checkbox rather than a genuine improvement tool. The last one is the most common reason this method fails in practice. If leadership reads the output and files it without funding any countermeasures, you have not done analysis, you have done paperwork. The effort is wasted regardless of how deep your chain goes. The technique itself is straightforward. Doing it well requires discipline at every step, honest verification of each answer, and the willingness to stop when the evidence runs out rather than filling gaps with assumptions. Most teams that attempt this skip verification and end up with a chain that looks logical but rests on a foundation of guesses. The difference between a useful analysis and a waste of time usually comes down to whether each link can be confirmed with something concrete before you move to the next why.
