Skip to content
KatafactsBeta

Problem solving / RCA

FMEA

FMEA is licensed CC BY 4.0. Attribution: Katafacts (katafacts.com).

Customise with AISkill file ↓

1 · What it is

What it is

A Failure Mode and Effects Analysis (FMEA) lists the specific ways a process or product could fail, then scores each one on three dimensions: severity (how bad the effect would be), occurrence (how likely the cause is), and detection (how likely current controls are to catch it before it matters) — each 1 to 10. Multiplying the three gives a risk priority number (RPN), the standard FMEA calculation, ranging from 1 to 1,000. The RPN isn't a precise probability; it's a ranking device that keeps attention on the failure modes that combine real severity, real likelihood, and a real detection gap, rather than whichever one someone happened to think of first.

2 · When to use it

When to use it — and when not to

Use it when

  • A process or product change is about to go live, and you want to think through what could go wrong before it does, not after.
  • You have several plausible failure modes and need a defensible way to rank which ones actually deserve attention first.
  • You want severity, occurrence, and detection scored explicitly and separately, not blended into one vague 'risk level' that hides which dimension is actually driving the concern.

Not when

  • A failure has already happened and you're investigating why — that's root cause analysis (an A3's 5-Whys), not a forward-looking risk analysis.
  • You're comparing a small set of options against multiple weighted criteria, not ranking failure modes — a weighted decision matrix is the right tool for that.
  • The severity, occurrence, or detection scores would just be guesses with no real basis — an FMEA built on invented scores ranks risks in the wrong order just as confidently as one built on real judgement.

3 · How to fill it in

How to fill it in

Process / product
What process, product, or system is this analysis covering?
Failure modes
For each item: what could go wrong, its effect, your severity score, the likely cause, your occurrence score, current controls, and your detection score. The risk priority number is computed from your three scores, never entered directly.
Narrative
Drafted from the items actually scored, naming the vital-few highest-risk items.

4 · What good looks like

What good looks like

The example below analyzes a newly-designed changeover procedure before it goes live — four failure modes scored on severity, occurrence, and detection, with the single highest-risk-priority-number item named as the vital-few risk to fix first, not just added to a long list.

Same example, as a downloadable xlsx workbook.

Download .xlsx

Failure Mode and Effects Analysis (FMEA)

FMEA — Line 2 paint booth changeover procedure

The new Line 2 paint booth changeover procedure, ahead of its kaizen-event rollout

Priya Nair, Line Supervisor · 2026-02-24

Process / product

The new Line 2 paint booth changeover procedure, ahead of its kaizen-event rollout

Failure modes

Top risk (vital few, by RPN): Nozzle cleaning step

Nozzle cleaning step

RPN 280

Failure mode: Cleaning cycle cut short under time pressure to hit the new 25-minute target

Effect: Color contamination carries into the next batch, caught only at final inspection

Cause: New target time creates pressure to skip the full cleaning cycle (S 8 · O 5 · D 7)

Recommended action: Add a timer-gated cleaning step the changeover can't proceed past until the full cycle completes (Marcus Webb 2026-02-27)

Color mixing during changeover

RPN 168

Failure mode: Wrong paint color mixed during a changeover

Effect: Full batch scrapped; changeover time doubles while the error is caught and corrected

Cause: Operator mixes from memory instead of the written recipe card (S 7 · O 4 · D 6)

Recommended action: Add a mandatory recipe-card check as the first step of every changeover, signed off before mixing starts (Marcus Webb 2026-02-27)

Queue board update

RPN 120

Failure mode: Board isn't updated in real time as small-batch orders arrive

Effect: Scheduler reverts to case-by-case judgement calls, defeating the sequencing rule the A3 put in place

Cause: No one is explicitly assigned to own keeping the board current (S 5 · O 6 · D 4)

Recommended action: Name the scheduler as the board's explicit owner, with a start-of-shift check as part of their standard work (Dana Ruiz 2026-03-02)

Fixture reuse across colors

RPN 90

Failure mode: A reused fixture carries residue from the previous color

Effect: Surface defect on the finished part

Cause: The new changeover procedure doesn't specify fixture-swap timing (S 6 · O 3 · D 5)

Recommended action: Add a fixture-swap checkpoint to the procedure, tied to color family rather than left to judgement (Marcus Webb 2026-03-06)

Narrative

The nozzle cleaning step is the single vital-few risk here, at a risk priority number well clear of everything else on the list — the new 25-minute target creates real pressure to shorten a cleaning cycle that has no independent check behind it, and a missed contamination event wouldn't be caught until final inspection. The color-mixing risk sits second, sharing the same root pattern (the redesigned procedure moved faster without adding the verification steps the old, slower procedure implicitly had time for), but it's the nozzle cleaning step that should get the recommended action actually built and verified first.

5 · Common mistakes

Common mistakes

  • Scoring severity, occurrence, and detection from a rough overall impression instead of thinking through each dimension separately.

    A failure mode that's severe but well-detected needs a different fix than one that's mild but invisible until it's too late — collapsing the three dimensions into one gut-feel number erases exactly the distinction that makes the risk priority number useful.

  • Treating every item on the list as equally worth fixing.

    Same vital-few discipline as every other artifact in this catalogue — a long list of scored risks with no ranking just moves the prioritization problem downstream instead of solving it.

  • Writing a recommended action that doesn't actually change severity, occurrence, or detection.

    "Communicate more clearly" or "be more careful" doesn't move any of the three numbers — a real recommended action targets one of them specifically (a physical control, a gate the process can't skip, a check that catches the failure earlier).

6 · What it connects to

What it connects to

upstream

  • A3 problem solving

    A failure an A3's root cause analysis already traced to a specific cause is exactly the kind of item an FMEA formalizes and ranks alongside other risks in the same process.

  • Scoping canvas

    A newly-scoped process or procedure change is what an FMEA should analyze before it rolls out, not after something's already gone wrong with it.

downstream

  • Countermeasure matrix

    When an FMEA surfaces more recommended actions than can be tackled at once, a countermeasure matrix is where they'd get ranked by impact and effort before committing.

7 · Where AI helps

Where AI helps

Judgement — stays yours

  • Deciding the real severity, occurrence, and detection scores for each failure mode
  • Deciding whether a recommended action actually targets the risk it's linked to

Analysis — AI helps

  • Tightening raw failure-mode, effect, cause, and controls notes into specific, scannable language
  • Drafting the narrative naming the vital-few highest-risk items
  • Drafting a recommended action grounded in the notes given, never inventing a fix with no basis in them

Drudgery — automated

  • Computing the risk priority number from severity, occurrence, and detection, every time a score changes
  • Ranking items by risk priority number
  • Exporting to xlsx in the house format

9 · Rate this kata

Rate this kata