The Process That Keeps You Up at 2 AM
ISO 14971 is the spine of medical device risk management, but knowing the standard exists is very different from actually living through it during an audit. The standard demands that you identify hazards, estimate risks, evaluate them against acceptability criteria, and control them. Then repeat that cycle for the life of the device. It sounds mechanical because it is mechanical. The problem is that the mechanical parts rarely work cleanly in practice. You start with a risk analysis. This is where you take a device apart conceptually and list every way it could cause harm. Electrical shock, biological contamination, software failure, user error, material degradation, environmental stress during shipping — the list goes on. You document each hazard, its associated hazardous situation, and the possible harm. A simple glucose monitor might have a hazard like "incorrect test strip calibration" leading to a "misdiagnosis" harm. A ventilator has a hazard around "pressure regulator failure" leading to "barotrauma to patient lungs." The analysis is only as good as the breadth of what you consider. After the analysis comes risk estimation. You assign likelihood and severity to each identified risk. Severity gets scored on a scale — typically 1 to 4 or 1 to 5 depending on your organization. Likelihood follows the same pattern. You multiply them together or reference a risk matrix to land on a risk level. This step is where most people get sloppy. Severity feels arbitrary until you're defending a score of 3 versus a 4 to an auditor who decided that night they were going to find something. I have seen teams use the same severity score for a blister rash and a permanent tissue injury because they were tired and the scoring table looked similar enough. It does not work that way. Severity is tied to the worst credible outcome, not the most likely one. Document your rationale for every score, or you will be rewriting it during a notification response.
Risk evaluation comes next. You compare each estimated risk against your organization's acceptability criteria. Risks that are acceptable get documented and moved on. Risks that are not acceptable require controls. This is the gate that separates the people who understand the standard from the people who just complete the paperwork. Acceptability criteria must be defined before you start. If you do not have pre-defined thresholds, you end up justifying every risk as acceptable through circular reasoning, which does not hold up under scrutiny.
Controls and the Downward Spiral
When a risk is not acceptable, you apply controls. The hierarchy goes from most effective to least: inherent safety by design, protective measures in the device itself, and information for safety — warnings, labels, instructions. This hierarchy matters because auditors check whether you followed it. I once had a device where the entire risk file was built around a warning label for an electrical hazard that could be eliminated by a simple enclosure redesign. The warning was cheaper and faster. The risk file looked compliant on the surface. During a Notified Body audit, the assessor asked me why we did not redesign the enclosure. I could not give her a technical answer, only a business one. She was not impressed. We went back, redesigned it, and rewrote approximately forty pages of documentation. It took three weeks. I still think about that every time someone suggests cutting a control down to a warning. Residual risk after controls must be evaluated again. Controls can introduce new hazards. A filter that catches particulates might increase pressure drop and create a new failure mode. A software update that patches a vulnerability might change the validation state and trigger a whole new round of testing. Every control you add needs its own risk assessment. This is the part that slows projects down. It is also the part that most teams underdocument.
Get the Full Details

Where Beginners Go Wrong
The biggest mistake I see is treating risk management as a product development phase rather than a lifecycle process. You produce a file, submit it, and then ignore it until the next audit or incident. ISO 14971 explicitly requires ongoing monitoring. Post-market feedback, field complaints, complaint handling, adverse event reporting — all of this feeds back into the risk file. When a customer reports a intermittent connectivity issue with your infusion pump, that data needs to trigger a review of the associated risk analysis. Most companies do not do this systematically. They collect the complaint, close the ticket, and move on. The risk file becomes a static artifact instead of a living document. This is a compliance gap and a patient safety gap. Another common failure is over-reliance on FMEA without context. Failure Mode and Effects Analysis is useful, but it is a tool, not a methodology. You can run a perfect FMEA on a component that was never identified as a hazard in the first place. The hazard identification step must come first and it must be thorough. I worked on a wearable patch device where our initial FMEA focused entirely on the adhesive and skin contact surfaces. We completely missed the battery management system as a hazard source because the system architecture diagram we were using did not show the BMS as a separate block. It was hidden inside the power module symbol. When a customer returned a unit with swelling, our risk file had nothing to reference. We had to rebuild the entire hazard identification from scratch. That experience changed how I approach system architecture diagrams permanently. Every block, even ones that look like black boxes, needs to be unpacked during hazard analysis.
What Nobody Tells You About Benign Risks
There is a category of risk that tends to get short-changed in medium-complexity devices: the benign-but-realistic scenario. These are situations where the harm is minor but the probability is high, and they accumulate into significant aggregate risk. A device that causes temporary skin irritation from its adhesive has a low individual severity score. But if ten thousand units are shipped and five hundred users report irritation, that becomes a regulatory and reputational problem. The standard requires you to consider the total risk posed by the device, not just the risk from a single use. I usually add a specific section in my risk files that tracks high-frequency low-severity scenarios separately. It takes extra time during the analysis phase, maybe twenty minutes per device, but it prevents embarrassing moments when post-market data shows a pattern that was invisible in the original assessment. Software risk is another area where people systematically underestimate. IEC 62304 intersects with ISO 14971 in ways that are often glossed over. A software bug that causes a minor display error in a diagnostic tool might seem harmless until you connect it to a clinical decision pathway. The risk is not in the software failure itself, it is in what the clinician does based on the incorrect information. Your risk analysis needs to trace software faults through their impact on the full clinical workflow. Most teams stop the trace at the device boundary. That is a gap.
A Practical Workflow That Does Not Waste Time
Start with the intended use statement. If this is vague or incomplete, everything downstream is compromised. I have seen risk files where the intended use was written as "for clinical monitoring" and the entire analysis was built on assumptions rather than facts. Define the indication, the population, the environment of use, and the user profile. A home-use device has a completely different risk landscape than an ICU device. The same hardware in different environments generates different hazards. Vibration, temperature, humidity, power quality — these are not trivial environmental factors. They are hazard sources. Use a combination of hazard identification techniques rather than relying on a single method. Brainstorming is good for getting the broad list. Checklists help catch common categories. Fault tree analysis is useful for complex systems where you need to understand how multiple failures interact. I normally run brainstorming sessions with cross-functional input — engineering, clinical, regulatory, and manufacturing. Manufacturing is the one people forget. They know which components fail in the field, which assembly steps introduce variability, and which materials degrade. Their input reduces the chance that you miss a hazard that exists only in production, not in design. Document assumptions explicitly. When you assume a certain failure rate from a supplier's data sheet, write that down. When you assume a user will follow a warning, note that assumption. These assumptions become vulnerabilities when post-market data contradicts them. An auditor can pick apart undocumented assumptions much faster than they can critique a documented and justified one. Justification is cheap insurance.

Risk control verification and validation are often conflated. Verification asks whether the control works as designed. Validation asks whether the control works in practice. A software alarm that triggers at 99 percent of simulated fault conditions is verified. A software alarm that a trained nurse misses because the sound frequency blends with background ICU noise is not validated. I now include a validation step for every high-severity control, usually through usability testing or simulated clinical environment testing. It adds about two weeks to the timeline for each major control, but it eliminates the rework that happens when an audit reveals that your controls are theoretical rather than practical.
Real Problems and What I Do About Them
Here is a specific edge case that caught me off guard. We were managing risk for a portable imaging device. The hazard analysis covered electrical safety, mechanical integrity, software reliability, and electromagnetic compatibility. The risk file was clean. Then a user in a rural clinic reported that the device produced image artifacts when a mobile phone was within one meter. We had EMC testing that passed the relevant standards. The test protocol did not include a realistic usage scenario with multiple active devices in close proximity. The risk was real, it was not acceptable under our criteria, and our risk file had no entry for it because it fell outside the test assumptions. The workaround was to add a usage scenario hazard category to our standard analysis template. Instead of only listing component-level and system-level hazards, we now include a category for contextual interference — environmental and operational factors that are outside the device's direct control but affect its performance in real use. This is not required by the letter of ISO 14971, but it is required by good practice. It added about an hour of work per device during the hazard identification phase and prevented a serious compliance gap on that project and several that followed. Another issue that comes up constantly is the treatment of legacy data. When you are working with an existing device and updating its risk file for a revision, you inherit risk analyses from previous versions. Some of that data may be outdated — older failure rates, superseded standards, obsolete components. I treat legacy risk data as presumption rather than proof. Every legacy entry gets reviewed against current standards, current supplier data, and current clinical understanding. This review typically takes one to two days for a moderately complex device and eliminates the risk of carrying forward something that is no longer valid. Skipping this review is one of the fastest ways to create a notification-ready situation.
The Limitations You Should Expect
Risk management is not a substitute for quality systems. A perfect risk file cannot compensate for poor manufacturing control. Conversely, excellent manufacturing control cannot fully compensate for a risk file that missed a fundamental hazard. The two systems are interdependent and both must function. If your design history file and your risk file diverge — if the design changes are not reflected in the risk analysis — you have a structural compliance problem that no amount of paperwork will fix. Quantification has hard limits. Severity scores are subjective by nature. Likelihood estimates derived from historical data may not apply to new designs or new populations. Risk matrices reduce complex probability distributions to colored boxes, which makes communication easier and accuracy harder. I use quantitative failure data wherever it exists, but I acknowledge the uncertainty when it does not. The risk file should reflect this honestly rather than pretending the numbers are more precise than they are. Auditors who understand the field respect honest uncertainty more than false precision. The standard assumes rational actors making rational decisions. It does not account well for systemic organizational failures — rushed timelines, understaffed review processes, pressure to ship. These are the conditions that produce the risk files I described earlier, the ones that look complete on the surface and fall apart under scrutiny. The best defense is institutional discipline: mandatory cross-functional review gates, independent verification of risk decisions, and a culture where raising a risk concern is treated as a contribution rather than an obstruction. Without those cultural elements, the process becomes a compliance exercise rather than a safety mechanism.
