Skip to content
KatafactsBeta

Problem solving / root cause analysis

Failure mode and effects analysis (FMEA)

Failure mode and effects analysis (FMEA) is licensed CC BY 4.0. Attribution: Katafacts (katafacts.com).

Customise with AISkill file ↓

1 · What it is

What it is

A failure mode and effects analysis (FMEA) lists the specific ways a process or product could fail, then scores each one on three dimensions: severity (how bad the effect would be), occurrence (how likely the cause is), and detection (how likely current controls are to catch it before it matters) — each 1 to 10. Severity comes first: a failure mode scored 9 or 10 on severity (someone could get hurt, or a legal or regulatory requirement is broken) needs a recommended action, or a written reason why the current controls are already enough, whatever its other scores. Multiplying the three scores gives a risk priority number (RPN), from 1 to 1,000, and this tool shows it as a secondary ranking for everything else. Be careful with it: very different risks can share the same RPN, and a severe failure that's rare and well detected can score lower than a trivial one that happens often. That's why the 2019 FMEA handbook from the Automotive Industry Action Group (AIAG) and the Verband der Automobilindustrie (VDA), the German automotive industry association, stopped ranking by RPN and replaced it with Action Priority: a lookup table that rates each failure mode high, medium, or low from its severity, occurrence, and detection scores, weighting severity most, then occurrence, then detection. If your customer or industry works to that handbook, use its Action Priority table to decide what to act on. Sources: AIAG & VDA FMEA Handbook, first edition (2019); AIAG Potential FMEA reference manual, fourth edition (2008), which already advised against using an RPN threshold to decide whether action is needed.

2 · When to use it

When to use it — and when not to

Use it when

  • A process or product change is about to go live, and you want to think through what could go wrong before it does, not after.
  • You have several plausible failure modes and need a defensible way to rank which ones actually deserve attention first.
  • You want severity, occurrence, and detection scored explicitly and separately, not blended into one vague 'risk level' that hides which dimension is actually driving the concern.

Not when

  • A failure has already happened and you're investigating why — that's root cause analysis (an A3's 5-Whys), not a forward-looking risk analysis.
  • You're comparing a small set of options against multiple weighted criteria, not ranking failure modes — a weighted decision matrix is the right tool for that.
  • The severity, occurrence, or detection scores would just be guesses with no real basis — an FMEA built on invented scores ranks risks in the wrong order just as confidently as one built on real judgement.

3 · How to fill it in

How to fill it in

Process / product
What process, product, or system is this analysis covering?
Failure modes
For each item: what could go wrong, its effect, your severity score, the likely cause, your occurrence score, current controls, and your detection score. The risk priority number is computed from your three scores, never entered directly.
Narrative
Drafted from the items actually scored: any severity 9 or 10 item first, then the vital-few highest risk priority number items.

4 · What good looks like

What good looks like

The example below analyzes a newly-designed changeover procedure before it goes live — four failure modes scored on severity, occurrence, and detection. It checks severity first (nothing scores 9 or 10, so nothing jumps the queue on severity alone), then names the single highest risk priority number item as the vital-few risk to fix first, not just added to a long list.

Same example, as a downloadable xlsx workbook.

Failure Mode and Effects Analysis (FMEA)

FMEA — Line 2 paint booth changeover procedure

The new Line 2 paint booth changeover procedure, ahead of its kaizen-event rollout

Priya Nair, Line Supervisor · 2026-02-24

Process / product

The new Line 2 paint booth changeover procedure, ahead of its kaizen-event rollout

Failure modes

Top risk (vital few, by RPN): Nozzle cleaning step

Nozzle cleaning step

RPN 280

Failure mode: Cleaning cycle cut short under time pressure to hit the new 25-minute target

Effect: Color contamination carries into the next batch, caught only at final inspection

Cause: New target time creates pressure to skip the full cleaning cycle (S 8 · O 5 · D 7)

Recommended action: Add a timer-gated cleaning step the changeover can't proceed past until the full cycle completes (Marcus Webb — 2026-02-27)

Color mixing during changeover

RPN 168

Failure mode: Wrong paint color mixed during a changeover

Effect: Full batch scrapped; changeover time doubles while the error is caught and corrected

Cause: Operator mixes from memory instead of the written recipe card (S 7 · O 4 · D 6)

Recommended action: Add a mandatory recipe-card check as the first step of every changeover, signed off before mixing starts (Marcus Webb — 2026-02-27)

Queue board update

RPN 120

Failure mode: Board isn't updated in real time as small-batch orders arrive

Effect: Scheduler reverts to case-by-case judgement calls, defeating the sequencing rule the A3 put in place

Cause: No one is explicitly assigned to own keeping the board current (S 5 · O 6 · D 4)

Recommended action: Name the scheduler as the board's explicit owner, with a start-of-shift check as part of their standard work (Dana Ruiz — 2026-03-02)

Fixture reuse across colors

RPN 90

Failure mode: A reused fixture carries residue from the previous color

Effect: Surface defect on the finished part

Cause: The new changeover procedure doesn't specify fixture-swap timing (S 6 · O 3 · D 5)

Recommended action: Add a fixture-swap checkpoint to the procedure, tied to color family rather than left to judgement (Marcus Webb — 2026-03-06)

Narrative

Severity first: no failure mode here scores 9 or 10 (nothing that could hurt someone or break a legal requirement), so nothing needs action on severity alone and the rest can be ranked by risk priority number. The nozzle cleaning step is the single vital-few risk here, at a risk priority number well clear of everything else on the list — the new 25-minute target creates real pressure to shorten a cleaning cycle that has no independent check behind it, and a missed contamination event wouldn't be caught until final inspection. The color-mixing risk sits second, sharing the same root pattern (the redesigned procedure moved faster without adding the verification steps the old, slower procedure implicitly had time for), but it's the nozzle cleaning step that should get the recommended action actually built and verified first.

5 · Common mistakes

Common mistakes

  • Ranking purely by risk priority number and ignoring a high-severity failure mode because its number is low.

    A failure that could hurt someone but rarely happens can score lower than a nuisance that happens every week. Severity 9 or 10 gets an action, or a written reason the controls are enough, regardless of the number — the same reason the 2019 handbook's Action Priority table weights severity most.

  • Setting a risk priority number cut-off, such as "act on anything over 100".

    The number isn't a true measure of risk — the scales aren't evenly spaced and many different score combinations give the same product. A fixed cut-off invites nudging a score down by one to get under the line instead of reducing the risk.

  • Scoring severity, occurrence, and detection from a rough overall impression instead of thinking through each dimension separately.

    A failure mode that's severe but well-detected needs a different fix than one that's mild but invisible until it's too late — collapsing the three dimensions into one gut-feel number erases exactly the distinction that makes the risk priority number useful.

  • Treating every item on the list as equally worth fixing.

    Same vital-few discipline as every other artifact in this catalogue — a long list of scored risks with no ranking just moves the prioritization problem downstream instead of solving it.

  • Writing a recommended action that doesn't actually change severity, occurrence, or detection.

    "Communicate more clearly" or "be more careful" doesn't move any of the three numbers — a real recommended action targets one of them specifically (a physical control, a gate the process can't skip, a check that catches the failure earlier).

6 · What it connects to

What it connects to

upstream

  • A3 problem solving

    A failure an A3's root cause analysis already traced to a specific cause is exactly the kind of item an FMEA formalizes and ranks alongside other risks in the same process.

  • Scoping canvas

    A newly-scoped process or procedure change is what an FMEA should analyze before it rolls out, not after something's already gone wrong with it.

downstream

  • Countermeasure matrix (not in the catalogue yet)

    When an FMEA surfaces more recommended actions than can be tackled at once, a countermeasure matrix is where they'd get ranked by impact and effort before committing.

Part of these playbooks

7 · Where AI helps

Where AI helps

Judgement — stays yours

  • Deciding the real severity, occurrence, and detection scores for each failure mode
  • Deciding whether a recommended action actually targets the risk it's linked to

Analysis — AI helps

  • Tightening raw failure-mode, effect, cause, and controls notes into specific, scannable language
  • Drafting the narrative naming the vital-few highest-risk items
  • Drafting a recommended action grounded in the notes given, never inventing a fix with no basis in them

Drudgery — automated

  • Computing the risk priority number from severity, occurrence, and detection, every time a score changes
  • Ranking items by risk priority number
  • Exporting to xlsx in the house format

9 · Rate this kata

Rate this kata