Problem solving / RCA
FMEA
1 · What it is
What it is
A Failure Mode and Effects Analysis (FMEA) lists the specific ways a process or product could fail, then scores each one on three dimensions: severity (how bad the effect would be), occurrence (how likely the cause is), and detection (how likely current controls are to catch it before it matters) — each 1 to 10. Multiplying the three gives a risk priority number (RPN), the standard FMEA calculation, ranging from 1 to 1,000. The RPN isn't a precise probability; it's a ranking device that keeps attention on the failure modes that combine real severity, real likelihood, and a real detection gap, rather than whichever one someone happened to think of first.
2 · When to use it
When to use it — and when not to
Use it when
- A process or product change is about to go live, and you want to think through what could go wrong before it does, not after.
- You have several plausible failure modes and need a defensible way to rank which ones actually deserve attention first.
- You want severity, occurrence, and detection scored explicitly and separately, not blended into one vague 'risk level' that hides which dimension is actually driving the concern.
Not when
- A failure has already happened and you're investigating why — that's root cause analysis (an A3's 5-Whys), not a forward-looking risk analysis.
- You're comparing a small set of options against multiple weighted criteria, not ranking failure modes — a weighted decision matrix is the right tool for that.
- The severity, occurrence, or detection scores would just be guesses with no real basis — an FMEA built on invented scores ranks risks in the wrong order just as confidently as one built on real judgement.
3 · How to fill it in
How to fill it in
- Process / product
- What process, product, or system is this analysis covering?
- Failure modes
- For each item: what could go wrong, its effect, your severity score, the likely cause, your occurrence score, current controls, and your detection score. The risk priority number is computed from your three scores, never entered directly.
- Narrative
- Drafted from the items actually scored, naming the vital-few highest-risk items.
4 · What good looks like
What good looks like
The example below analyzes a newly-designed changeover procedure before it goes live — four failure modes scored on severity, occurrence, and detection, with the single highest-risk-priority-number item named as the vital-few risk to fix first, not just added to a long list.
Same example, as a downloadable xlsx workbook.
Download .xlsxFailure Mode and Effects Analysis (FMEA)
FMEA — Line 2 paint booth changeover procedure
The new Line 2 paint booth changeover procedure, ahead of its kaizen-event rollout
Priya Nair, Line Supervisor · 2026-02-24
Process / product
The new Line 2 paint booth changeover procedure, ahead of its kaizen-event rollout
Failure modes
Top risk (vital few, by RPN): Nozzle cleaning step
Nozzle cleaning step
RPN 280Failure mode: Cleaning cycle cut short under time pressure to hit the new 25-minute target
Effect: Color contamination carries into the next batch, caught only at final inspection
Cause: New target time creates pressure to skip the full cleaning cycle (S 8 · O 5 · D 7)
Recommended action: Add a timer-gated cleaning step the changeover can't proceed past until the full cycle completes (Marcus Webb — 2026-02-27)
Color mixing during changeover
RPN 168Failure mode: Wrong paint color mixed during a changeover
Effect: Full batch scrapped; changeover time doubles while the error is caught and corrected
Cause: Operator mixes from memory instead of the written recipe card (S 7 · O 4 · D 6)
Recommended action: Add a mandatory recipe-card check as the first step of every changeover, signed off before mixing starts (Marcus Webb — 2026-02-27)
Queue board update
RPN 120Failure mode: Board isn't updated in real time as small-batch orders arrive
Effect: Scheduler reverts to case-by-case judgement calls, defeating the sequencing rule the A3 put in place
Cause: No one is explicitly assigned to own keeping the board current (S 5 · O 6 · D 4)
Recommended action: Name the scheduler as the board's explicit owner, with a start-of-shift check as part of their standard work (Dana Ruiz — 2026-03-02)
Fixture reuse across colors
RPN 90Failure mode: A reused fixture carries residue from the previous color
Effect: Surface defect on the finished part
Cause: The new changeover procedure doesn't specify fixture-swap timing (S 6 · O 3 · D 5)
Recommended action: Add a fixture-swap checkpoint to the procedure, tied to color family rather than left to judgement (Marcus Webb — 2026-03-06)
Narrative
The nozzle cleaning step is the single vital-few risk here, at a risk priority number well clear of everything else on the list — the new 25-minute target creates real pressure to shorten a cleaning cycle that has no independent check behind it, and a missed contamination event wouldn't be caught until final inspection. The color-mixing risk sits second, sharing the same root pattern (the redesigned procedure moved faster without adding the verification steps the old, slower procedure implicitly had time for), but it's the nozzle cleaning step that should get the recommended action actually built and verified first.
5 · Common mistakes
Common mistakes
Scoring severity, occurrence, and detection from a rough overall impression instead of thinking through each dimension separately.
A failure mode that's severe but well-detected needs a different fix than one that's mild but invisible until it's too late — collapsing the three dimensions into one gut-feel number erases exactly the distinction that makes the risk priority number useful.
Treating every item on the list as equally worth fixing.
Same vital-few discipline as every other artifact in this catalogue — a long list of scored risks with no ranking just moves the prioritization problem downstream instead of solving it.
Writing a recommended action that doesn't actually change severity, occurrence, or detection.
"Communicate more clearly" or "be more careful" doesn't move any of the three numbers — a real recommended action targets one of them specifically (a physical control, a gate the process can't skip, a check that catches the failure earlier).
6 · What it connects to
What it connects to
upstream
A3 problem solving
A failure an A3's root cause analysis already traced to a specific cause is exactly the kind of item an FMEA formalizes and ranks alongside other risks in the same process.
Scoping canvas
A newly-scoped process or procedure change is what an FMEA should analyze before it rolls out, not after something's already gone wrong with it.
downstream
Countermeasure matrix
When an FMEA surfaces more recommended actions than can be tackled at once, a countermeasure matrix is where they'd get ranked by impact and effort before committing.
7 · Where AI helps
Where AI helps
Judgement — stays yours
- Deciding the real severity, occurrence, and detection scores for each failure mode
- Deciding whether a recommended action actually targets the risk it's linked to
Analysis — AI helps
- Tightening raw failure-mode, effect, cause, and controls notes into specific, scannable language
- Drafting the narrative naming the vital-few highest-risk items
- Drafting a recommended action grounded in the notes given, never inventing a fix with no basis in them
Drudgery — automated
- Computing the risk priority number from severity, occurrence, and detection, every time a score changes
- Ranking items by risk priority number
- Exporting to xlsx in the house format
9 · Rate this kata
