🎯 Quick Answer
What: An FMEA (Failure Mode and Effects Analysis) is a structured list of the ways a process could fail, written before the failures happen. Each row names one step, one way it could fail, and three scores: severity (how bad the effect would be), occurrence (how often the cause is likely), and detection (how likely current controls are to catch it). Multiplied, the scores give a risk priority number that ranks the rows.
Why it matters: The people closest to the work already know most of an operation's next failures. Without a written, scored list, that knowledge never meets a decision, and each failure arrives as a surprise to everyone except the people who predicted it.
How to apply: Walk the process one step at a time with the people who run it. For each step, ask how it could fail, then score the effect, the cause and the controls against scales agreed in advance. Act on the highest-ranked rows and on any row with a severe effect, prefer actions that prevent the cause over actions that improve detection, and re-score after every fix.
The payoff: Failures that were already predictable get prevented instead of explained, and attention goes to the failure that deserves it first rather than the one mentioned most recently.
Ask any crew at shift change what is going to go wrong next, and you will have an answer in under a minute.
The bearing that has started to sing. The supplier whose deliveries arrive a little later every month. The new step on the night shift that only one person really knows how to do. The people closest to the work carry a remarkably accurate forecast of the operation's next failures. It lives in their heads, in the margin of a maintenance log, in a comment at a handover that nobody wrote down.
Almost none of it is written anywhere it could be used. So the forecast never meets a decision, and when the failure arrives it is a surprise to everyone except the people who predicted it.
An FMEA is that forecast, written down in advance. This note is about what it is, how to fill one in honestly, and the handful of rules that separate a useful afternoon from a spreadsheet of guesses.
What an FMEA is, and what it is not
Most analysis in a Lean Six Sigma project works after the fact. A fishbone and a Five Whys both begin with something that has already gone wrong and work backward to why it happened. An FMEA runs the other direction. It begins with a failure that has not happened yet and asks how it could.
It is not a brainstorm, although it starts like one. A cause-and-effect session collects every plausible reason for one problem. An FMEA walks a whole process, step by step, and asks the same question of each step: how could this step fail to do its job? Each answer becomes a row, and each row is scored, so the finished document is not a list of worries but a ranking of them.
One row at a time
A single row is the whole method in miniature. Take one step from a press shop: setting up the press for the next job.
The failure mode, in plain words: the previous job's settings carry over. Name the failure, not the cause; "the settings carry over" is what goes wrong, and why it happens comes later in the row.
The effect and its severity. Slightly out-of-spec parts move on to assembly and fail there, or worse, do not fail there. Severity asks how bad the effect is for whoever receives the output, on a scale of 1 to 10. The team scores it 8.
The cause and its occurrence. The setup sheet is not updated when the job changes. Occurrence asks how often that cause is likely to happen. A few times a year: 3.
The controls and detection. Today, a first-piece check, when someone remembers to do it. Detection asks how likely current controls are to catch the failure before it leaves, and it runs backward: 1 means almost certain to be caught, 10 means nothing catches it. The team scores it 6.
The risk priority number. 8 x 3 x 6 = 144. On its own, 144 means nothing. Its only job is to sort this row against every other row in the document.
The action. Load the settings from the job order itself, so the old ones cannot carry over. Then score the row again.
A finished FMEA is dozens of rows like that one, and the ranking of their numbers is where the conversation about what to do first begins.
Reading the scores honestly
The arithmetic is trivial. The honesty is not, and four habits decide whether the ranking can be trusted.
Anchor the scales before the first row. Agree what a 3, a 7 and a 10 mean for each score, in the operation's own terms, before anything is rated. Without anchors, a 7 on row 3 and a 7 on row 30 mean different things, and the scores drift toward whatever the most confident voice in the room feels.
Treat the number as a sorting device, not a grade. A risk priority number says where to look first. It never says the rows beneath are safe. Ranking failures by risk is also a different act from ranking improvement projects by value; the first decides where attention goes inside a process, the second decides which work gets funded at all.
Let severity speak on its own. Multiplication can hide a catastrophe. A failure severe enough to reach a customer or stop the operation earns an action even when its occurrence is low and its combined score is modest. The numbers sort. They do not excuse.
Distrust flattering detection scores. Of the three, detection is the score teams overrate most, because a careful colleague feels like a control. Ask who catches the failure, with what, at which step, and whether it is caught every time or only when someone happens to be paying attention. If the honest answer is "someone might notice", score it as if nothing does.
Prevent, then detect
When it is time to act, the order of preference matters. An action that stops the cause from occurring beats one that improves the odds of catching the failure afterward. Inspection bolted onto a weak step changes the odds of catching the failure and does nothing to the odds of the failure itself.
Then score the row again. An action has not reduced a risk until the row has been re-scored on evidence and the number has actually dropped. The re-scored document is also the record that shows the work was worth doing, which is the same verification habit that separates a tested cause from a believed one.
When to run one, and who should be in the room
In the DMAIC arc, an FMEA does its first job in Analyze, where it ranks the ways the current process fails so the project's attention lands in the right place. It often returns in Improve, pointed at the new design, to stress-test a solution before the pilot. Outside projects, the moments that most deserve one are the moments of change: a new line, a new supplier, a changeover, a system cutover, a process being handed to a new team.
The people who run the step belong in the room. The engineer knows the design; the operator knows the Tuesday-night version of it, and the failures worth finding usually live in the second one. Give the room's pessimist the floor, write every answer down before anyone argues with it, and keep naming a risk separate from owning it, so saying it out loud costs nothing.
Finally, keep the document alive. An FMEA filed when the project closes is a record of what a team once worried about. One reopened whenever the process changes is how an operation stays ahead of its next failure.
I write one of these every week, drawn from what actually happens on projects rather than what the methodology says should. You can join the newsletter here.
FAQ
What does FMEA stand for?
Failure Mode and Effects Analysis. A failure mode is a way a process step could fail to do its job; an effect is what that failure would do to whoever receives the output. The analysis lists the failure modes of a process, scores each one, and ranks them so the riskiest get action first. A process FMEA looks at the steps of a process; a design FMEA applies the same logic to the parts of a product.
How is the risk priority number calculated?
Multiply the three scores: severity x occurrence x detection, each on a 1 to 10 scale, for a number between 1 and 1,000. Severity rates the effect, occurrence rates how often the cause is likely, and detection rates how likely current controls are to catch the failure, with 10 meaning nothing would. The number is only meaningful relative to the other rows in the same document, scored against the same anchored scales.
Is there an RPN threshold that requires action?
There is no universal cutoff, and a fixed threshold invites people to score rows just under it. The sounder practice is to act on the highest-ranked rows first and on any row with a high severity score regardless of its total. Some industry handbooks, the automotive AIAG and VDA handbook among them, have replaced the single multiplied number with action-priority tables that weight severity first, for exactly that reason.
What is the difference between an FMEA and root cause analysis?
Timing and direction. Root cause analysis starts from a failure that has already happened and works backward to why it happened. An FMEA starts from failures that have not happened yet and works forward from each process step to how it could fail, how badly, and how likely it is to be caught. The two meet in practice: an FMEA's cause column often borrows candidates from past root cause work.
How often should an FMEA be updated?
Whenever the process changes, and whenever a failure happens that the document did not predict. A new supplier, a new product, a new shift pattern or an equipment change should all reopen it. A failure that was not on the list is the most useful update of all: it shows exactly where the team's forecast was blind.
A failure written down in advance can still be prevented. One discovered afterward can only be explained.









