Running a Proper FMEA Without Losing Your Mind
The process itself is simple on paper, which is part of why it gets botched so often. You list every component or step, identify how each one could fail, figure out what happens when it fails, rate the severity, the likelihood, and how well you can catch it, multiply those three numbers together, and rank from there. That is the standard AIAG-VDA approach most of us use. The problem is not the formula, it is the people filling it out. Severity, Occurrence, and Detection are each scored from one to ten, and the Risk Priority Number is just their product. A severe failure with high occurrence and poor detection gets a number over 200. Anything above 100 typically demands immediate corrective action in an automotive environment. Below 50, you might tolerate it depending on the application. Most teams miss the nuance here and treat RPN as gospel instead of a rough prioritization tool. I ran an Ejemplo De Analisis De Modos Y Efectos De Fallos for a powertrain control module last year where we had a solder joint failure on a passive component on the primary side of a DC-DC converter. Severity was an eight because the output went to zero with no warning. Occurrence scored a four based on IPC standards at the time, and detection was a nine since our automated test caught nothing. The RPN came out to 288, which should have been a red flag. But here is the thing that nobody tells you in the training sessions: the detection score is where most of your analysis collapses under its own weight.
We ended up redesigning the footprint, adding a second test point, and changing the component placement to reduce thermal stress during reflow. The detection score dropped to a three after we added in-circuit testing with boundary scan coverage. The RPN fell to 72. Took two weeks of work and three prototype revisions. Worth it.
How to Actually Build the Spreadsheet Without It Becoming Useless
Start with a boundary diagram. Map out every input and output, every power rail, every communication line. Then break the system down into sub-assemblies. Do not start scoring until every single failure mode for every single component is listed. I have seen teams skip this and go straight to scoring, which produces garbage numbers because they missed a failure mode entirely. The spreadsheet should have columns for item number, function, failure mode, effect, severity, cause, occurrence, current controls, detection, RPN, recommended actions, responsibility, target date, and results after implementation. Keep it in one workbook. Multiple files become a nightmare within six months.
The Counter-Intuitive Parts Nobody Talks About
First, severity is independent of your controls. Adding a sensor that detects a pressure drop does nothing to lower the severity of a seal failure. It only improves detection. Teams routinely downgrade severity scores because they have a mitigation in place. That is wrong. Severity is about the consequence to the customer or the system, not about your safety net.Get the Full Details

Second, occurrence and detection are not the same thing even though beginners treat them like twins. Occurrence is about how likely the root cause is to happen. Detection is about whether your current testing or inspection can find it before it reaches the next stage or the customer. A well-designed control system will have high occurrence but low detection if the failure mode is subtle. That combination is the most dangerous because the math looks better than it actually is. I encountered this on a hydraulic valve assembly where the solenoid coil had a moderate occurrence rate of four due to thermal cycling, but our electrical testing at 85 degrees Celsius detected the degradation mode poorly, scoring a seven for detection. The RPN was 63, which looked acceptable on paper. Two years later, field returns showed a twelve percent failure rate in cold climate applications. We had misclassified the detection score because our test was performed at room temperature, not at operating temperature. Fixing the test condition alone dropped detection to a two and revalidated the original risk assessment.
When FMEA Actually Fails and What to Use Instead
FMEA assumes you can enumerate every failure mode. That assumption breaks down for complex software systems, machine learning pipelines, or any system with emergent behavior. If your product has more than five interacting subsystems where the interaction itself can create novel failure states, FMEA will miss things. It always will. In those cases, combine it with fault tree analysis or try a system-level reliability block diagram first to identify interaction points, then run the FMEA on those specific interfaces rather than treating every component in isolation. The process also falls apart when the cross-functional team is not real. If one engineer fills out the document alone and sends it to manufacturing and quality for signatures, you have produced compliance paperwork, not a useful risk analysis. The whole point of FMEA is that the mechanic who assembles the part knows something the designer did not, and the field service technician knows something the tester did not. If those voices are not in the room during the session, the output will be inaccurate regardless of how carefully you fill out the scores. Update frequency matters more than most teams realize. An FMEA that is done once at the beginning of a product cycle and never touched again is worse than useless, because it creates false confidence. I recommend a quarterly review during active development and a mandatory update whenever a field return, a warranty claim, or a production scrap spike occurs that relates to the system you analyzed. Even if the change is minor, documenting it closes the loop and keeps the document living instead of archived.
There is no free download that will save you from doing this work properly. Templates exist everywhere, but they are starting points, not solutions. The value is in the discussion, not in the grid. Start with the template, fill it with your team, argue through the scores, and accept that the numbers will be wrong at first. They always are. The second and third iterations are where the real analysis happens.
