Root Cause Analysis – 5 Whys and FMEA
Learning Objectives
By the end of this lesson, the learner will be able to:
- Explain the purpose of Root Cause Analysis (RCA).
- Distinguish symptoms, contributing causes, and root causes.
- Explain the principles of the 5 Whys technique.
- Apply the 5 Whys method systematically.
- Recognize common weaknesses in 5 Whys analysis.
- Explain Failure Mode and Effects Analysis (FMEA).
- Distinguish failure mode, effect, cause, and control.
- Identify potential risks systematically using FMEA.
- Understand Severity, Occurrence, and Detection ratings.
- Explain the traditional Risk Priority Number (RPN) concept and its limitations.
- Understand Action Priority as an alternative risk-prioritization approach where applicable.
- Connect FMEA with Control Plans and process improvement.
- Develop appropriate preventive and detection controls.
- Use RCA and FMEA together to reduce recurrence of process problems.
1. Introduction
A Six Sigma project is not successful merely because a problem has been identified.
The Green Belt must determine why the problem occurs and develop actions that address the underlying causes.
Consider a manufacturing process with an increasing defect rate.
A superficial response might be:
“Inspect more products.”
Inspection may identify defective products, but it does not necessarily eliminate the reason the defects occur.
Root Cause Analysis seeks to move from:
Problem → Symptom → Cause → Root Cause → Corrective Action → Verification
Two important tools used by Green Belts are:
- 5 Whys
- Failure Mode and Effects Analysis (FMEA)
The 5 Whys is particularly useful for investigating an existing problem.
FMEA is particularly useful for systematically identifying and assessing potential failure modes and associated risks.
They serve different but complementary purposes.
2. Concept / Theory
2.1 What Is Root Cause Analysis?
Root Cause Analysis is a structured approach for identifying the underlying reason or reasons a problem occurs.
The purpose is not simply to explain what happened.
The purpose is to understand the causal chain sufficiently to determine where effective action can be taken.
A useful distinction is:
Symptom
The visible problem.
Example:
Customer orders are delivered late.
Immediate Cause
The direct condition associated with the problem.
Example:
Orders were released to production late.
Contributing Cause
A condition that helped create or worsen the problem.
Example:
Order-priority information was not consistently available.
Root Cause
An underlying cause that, when appropriately addressed, prevents or substantially reduces recurrence of the problem.
Example:
The order-management process has no standardized priority-classification rule integrated into order release.
The exact root cause must be established through evidence rather than simply selected because it sounds reasonable.
2.2 Why Root Cause Analysis Matters
Without effective RCA, organizations may repeatedly solve the same problem.
Common weak responses include:
- Retraining everyone.
- Reminding employees to be careful.
- Increasing inspection.
- Reworking defective products.
- Blaming an operator.
- Adjusting the process temporarily.
These actions may sometimes be appropriate, but they do not automatically address systemic causes.
An effective root-cause investigation asks:
What condition in the process allowed the problem to occur?
And then:
What evidence demonstrates that this condition contributes to the problem?
2.3 The 5 Whys
The 5 Whys is a simple questioning technique used to explore the causal chain behind a problem.
The team repeatedly asks:
Why did this happen?
The number five is not a strict requirement.
Sometimes three questions may be sufficient.
Sometimes more than five are necessary.
The objective is to reach a meaningful underlying cause supported by evidence.
2.4 Example of 5 Whys
Problem
A customer shipment was delivered late.
Why 1?
Why was the shipment late?
Because the order was not completed by the planned dispatch date.
Why 2?
Why was the order not completed by the planned dispatch date?
Because production received the order later than expected.
Why 3?
Why did production receive the order late?
Because the order-release process was delayed.
Why 4?
Why was order release delayed?
Because required order information was incomplete.
Why 5?
Why was required information incomplete?
Because the order-entry process did not include a standardized completeness check before release.
Potential Root Cause
The order-entry process lacked a standardized completeness-control mechanism before release.
The team should then verify this relationship using actual process data and records.
2.5 The 5 Whys Is Not a Blame Tool
A poor 5 Whys analysis can quickly become:
“Why did the operator make the mistake?”
→ “Because the operator was careless.”
This is generally too superficial.
The Green Belt should investigate the process conditions that allowed the error.
For example:
- Was the work instruction clear?
- Was the correct information available?
- Was the system designed to prevent the error?
- Was the equipment functioning correctly?
- Was training adequate?
- Was workload excessive?
- Was the process standard followed?
- Was the standard itself appropriate?
The objective is to improve the process, not merely identify an individual.
2.6 Evidence-Based Root Cause
A proposed root cause should be supported by evidence.
Useful evidence may include:
- Process data
- Time-series data
- Observation
- Transaction records
- Interviews
- Process maps
- Experimental evidence
- Stratification
- Pareto analysis
- Correlation or regression analysis
- Hypothesis testing
- FMEA
- Historical records
A hypothesis is not automatically a root cause.
The Green Belt should distinguish:
Possible Cause → Suspected Cause → Verified Cause
3. Key Topics
3.1 FMEA
Failure Mode and Effects Analysis (FMEA) is a structured method for identifying potential failure modes, understanding their effects and causes, evaluating risk, and determining actions to reduce risk.
FMEA can be applied:
- Before a process is implemented.
- During process design.
- During product development.
- During process improvement.
- After significant process changes.
- When recurring failures occur.
FMEA is therefore primarily a risk-prevention and risk-management tool.
3.2 Failure Mode
A failure mode describes how a process, product, or service could fail to meet its intended requirement.
Examples:
- Incorrect product quantity
- Wrong part installed
- Incorrect invoice amount
- Missing customer information
- Component diameter outside specification
- Incorrect medication dosage entry
A failure mode should be stated clearly enough that the team can evaluate its consequences and causes.
3.3 Effect
The effect describes the consequence of the failure mode.
For example:
Failure Mode: Incorrect product quantity
Effect: Customer receives an incomplete order.
Another possible effect could be:
Customer complaint and additional shipment required.
The effect should be considered from the perspective of the process, next operation, customer, safety, or other relevant stakeholder.
3.4 Cause
The cause describes why the failure mode might occur.
Example:
Failure Mode: Incorrect product quantity
Cause: Quantity entered incorrectly during order processing.
Potential underlying causes could include:
- Manual entry
- Ambiguous order format
- Poor interface design
- Lack of validation
- Incorrect master data
The team should avoid confusing the failure mode itself with its cause.
3.5 Current Controls
Controls are existing measures intended to:
- Prevent a failure from occurring.
- Detect a failure before it reaches the customer or next process.
Examples:
Prevention Control
Automated validation prevents an invalid quantity from being entered.
Detection Control
A final order audit identifies quantity discrepancies.
FMEA should document relevant existing controls because the risk assessment depends partly on the effectiveness of those controls.
3.6 Severity
Severity represents the significance of the effect if the failure occurs.
A higher severity rating generally represents a more serious consequence.
Examples of consequences that may warrant higher severity include:
- Safety impact
- Regulatory impact
- Major customer impact
- Significant financial impact
- Major operational disruption
Severity concerns the effect, not how frequently the failure occurs.
3.7 Occurrence
Occurrence represents how frequently or likely the cause/failure mode is expected to occur under the defined conditions.
Historical process data can provide useful evidence.
Examples include:
- Defect frequency
- Historical failure rates
- Complaint frequency
- Process data
- Field-failure history
Occurrence is therefore related to likelihood or frequency, not consequence severity.
3.8 Detection
Detection represents the effectiveness of the current controls in detecting the failure mode or its cause before the undesired effect reaches the next stage or customer, depending on the FMEA methodology being used.
A stronger detection capability generally corresponds to a lower detection risk.
Detection should not be confused with:
“How easy is the problem to see after it reaches the customer?”
The rating depends on the defined FMEA rating system and the timing/effectiveness of the control.
3.9 Traditional RPN
Traditional FMEA may calculate a Risk Priority Number (RPN):
RPN = Severity × Occurrence × Detection
For example:
- Severity = 8
- Occurrence = 5
- Detection = 4
Therefore:
RPN = 8 × 5 × 4 = 160
RPN can help prioritize attention, but it has limitations.
Different combinations of Severity, Occurrence, and Detection can produce the same RPN.
For example:
8 × 5 × 4 = 160
and
10 × 4 × 4 = 160
These combinations do not represent identical risk characteristics.
Therefore, RPN should not automatically be treated as the complete representation of risk.
Organizations should follow the FMEA methodology and risk-prioritization system applicable to their context.
3.10 Action Priority
Modern FMEA approaches may use Action Priority (AP) rather than relying solely on RPN.
Action Priority considers the combination of:
- Severity
- Occurrence
- Detection
and provides a structured way to prioritize recommended actions.
The important principle is:
Risk prioritization should consider the nature of the risk, not merely the multiplication of three numbers.
3.11 FMEA Types
Depending on the application, organizations may use different forms of FMEA, such as:
- Design FMEA (DFMEA)
- Process FMEA (PFMEA)
- System FMEA
- Service/process applications
For Green Belt process-improvement work, Process FMEA (PFMEA) is particularly relevant.
3.12 5 Whys vs FMEA
| Aspect | 5 Whys | FMEA |
|---|---|---|
| Primary purpose | Investigate causal chain | Identify and assess potential failures |
| Typical orientation | Existing problem | Potential and existing risks |
| Structure | Sequential questioning | Structured table |
| Main focus | Causes | Failure modes, effects, causes, controls |
| Risk rating | Usually not inherent | Severity, Occurrence, Detection / AP |
| Best use | Focused investigation | Systematic risk assessment |
| Output | Causal understanding | Prioritized risk-reduction actions |
They should not be treated as interchangeable tools.
4. Tools / Methodology
4.1 5 Whys Methodology
Step 1 — Define the Problem
State the problem clearly and specifically.
Weak:
“Quality is poor.”
Better:
“The first-pass yield for Product A fell from 96% to 89% during the last four weeks.”
Step 2 — Verify the Problem
Confirm that the problem actually exists using data.
Step 3 — Ask Why
Ask why the observed condition occurred.
Step 4 — Verify the Answer
Use evidence to determine whether the proposed cause is supported.
Step 5 — Continue
Ask why again until reaching a sufficiently actionable underlying cause.
Step 6 — Test the Cause
Where practical, use data or controlled testing to verify the causal relationship.
Step 7 — Develop Corrective Action
Address the verified cause.
Step 8 — Verify Effectiveness
Confirm that the problem has been reduced or eliminated.
4.2 5 Whys Quality Check
A good 5 Whys chain should:
- Begin with a clearly defined problem.
- Use evidence where available.
- Maintain logical causal connections.
- Avoid blaming individuals without examining system conditions.
- Avoid assumptions presented as facts.
- Identify actionable causes.
- Lead to a corrective or preventive action.
- Include verification of effectiveness.
4.3 FMEA Methodology
Step 1 — Define the Process
Identify the process being analyzed.
Step 2 — Break the Process into Steps
Use a process map or flowchart.
Step 3 — Identify Failure Modes
Ask:
“How could this process step fail to meet its intended requirement?”
Step 4 — Identify Effects
Ask:
“What happens if this failure occurs?”
Step 5 — Identify Causes
Ask:
“Why could this failure occur?”
Step 6 — Identify Current Controls
Determine how the failure is:
- Prevented
- Detected
Step 7 — Assess Risk
Apply the organization’s defined rating scales for:
- Severity
- Occurrence
- Detection
or the applicable Action Priority methodology.
Step 8 — Identify Recommended Actions
Potential actions include:
- Error-proofing
- Process redesign
- Automation
- Standardization
- Training
- Preventive maintenance
- Additional controls
- Supplier controls
- Improved measurement
- Improved detection
Step 9 — Assign Responsibility
Every significant action should have:
- Owner
- Target date
- Required resources
Step 10 — Reassess
After action implementation:
- Verify effectiveness.
- Update ratings where applicable.
- Update controls.
- Document residual risk.
5. Worked Example / Case Study
Case Study: Incorrect Invoice Amounts
A company receives customer complaints because some invoices contain incorrect amounts.
Step 1 — Define the Problem
Data show that invoice errors increased from:
0.8% to 2.4%
over the previous two months.
5.1 Applying 5 Whys
Problem
Incorrect invoice amount sent to customer.
Why 1?
Because the invoice contained an incorrect price.
Why 2?
Because the price entered into the transaction did not match the approved price.
Why 3?
Because the approved price was manually entered.
Why 4?
Why was the approved price manually entered?
Because the system did not automatically retrieve the approved price from the master-price database.
Why 5?
Why did the process depend on manual price entry?
Because the current process design did not include automated price validation before invoice release.
Potential Root Cause
The invoice process lacks an automated validation mechanism linking invoice prices to approved master-price data.
This proposed root cause should be verified using transaction records and process testing.
5.2 Applying FMEA
Consider the invoice-release process.
| Process Step | Failure Mode | Effect | Potential Cause | Current Control |
|---|---|---|---|---|
| Enter price | Incorrect price | Incorrect invoice | Manual entry | Operator review |
| Validate invoice | Error not detected | Customer receives incorrect invoice | Review may miss discrepancy | Sample audit |
| Release invoice | Incorrect invoice released | Complaint/rework | No automated price comparison | Final approval |
The team then evaluates:
- Severity
- Occurrence
- Detection
using the organization’s defined rating scales.
Suppose one failure mode receives:
- Severity = 8
- Occurrence = 5
- Detection = 4
Traditional RPN:
8 × 5 × 4 = 160
The team may prioritize the risk for improvement.
However, the team should also consider the individual rating values and the organization’s Action Priority methodology where applicable.
5.3 Improvement Action
The team introduces:
Automated comparison of invoice price against the approved master-price database before release.
This changes the control from primarily manual detection toward automated prevention/detection.
The team then verifies:
- Error rate
- System effectiveness
- False alarms
- Processing time
- Customer complaints
The FMEA and Control Plan should be updated after the validated process change.
6. Practical Application
Exercise 1 — 5 Whys
Problem:
A customer received an incorrect product.
Develop a 5 Whys analysis.
Consider possible process conditions such as:
- Order entry
- Product identification
- Picking
- Labeling
- Verification
- System controls
- Training
- Standard work
Do not automatically assign the cause to operator error.
Identify evidence needed to verify each proposed cause.
Exercise 2 — FMEA
Select one process and identify at least five failure modes.
For each failure mode document:
- Process step
- Failure mode
- Effect
- Potential cause
- Current prevention control
- Current detection control
- Severity
- Occurrence
- Detection
- Recommended action
- Action owner
- Completion date
Use the organization’s approved rating methodology.
Exercise 3 — Distinguishing Cause and Effect
For each statement, identify whether it is primarily a failure mode, effect, or cause.
A.
“Customer receives an incorrect invoice.”
Effect
B.
“Incorrect price entered.”
Failure mode
C.
“Approved price is manually entered without system validation.”
Potential cause
Discuss why clear terminology is important when constructing an FMEA.
Exercise 4 — Root Cause Verification
A team concludes:
“The operator was not careful.”
Ask:
- What evidence supports this conclusion?
- Was the work instruction clear?
- Was the correct information available?
- Was the process designed to prevent the error?
- Was the employee properly trained?
- Was the equipment functioning correctly?
- Did the error occur with other operators?
- Did workload or process conditions contribute?
- Can the proposed cause be tested?
The objective is to move from assumption to evidence.
6.1 RCA and FMEA Checklist
Before closing an RCA/FMEA activity, verify:
- Problem is clearly defined.
- Data confirm the problem.
- Symptoms are distinguished from causes.
- Potential causes are systematically identified.
- Evidence supports important causal conclusions.
- 5 Whys logic is coherent.
- Individual blame has not replaced process analysis.
- Failure modes are clearly stated.
- Effects are clearly identified.
- Causes are distinguished from effects.
- Existing controls are documented.
- Risk ratings follow the approved methodology.
- Recommended actions address meaningful causes.
- Action owners are assigned.
- Completion dates are defined.
- Effectiveness is verified.
- FMEA and Control Plan are updated where necessary.
7. Lesson Summary
Root Cause Analysis helps organizations move beyond symptoms and address the underlying conditions responsible for problems.
The major principles are:
- A symptom is not necessarily a root cause.
- Root causes should be supported by evidence.
- The 5 Whys explores a causal chain by repeatedly asking why.
- Five is not a mandatory number of questions.
- The 5 Whys should not become a blame exercise.
- FMEA systematically identifies potential failure modes and their effects and causes.
- FMEA documents existing prevention and detection controls.
- Severity describes the consequence of failure.
- Occurrence describes likelihood or frequency.
- Detection concerns the effectiveness of controls in identifying the failure or cause.
- Traditional RPN is calculated as Severity × Occurrence × Detection.
- RPN has limitations and should not be treated as the only basis for risk decisions.
- Action Priority may be used in modern FMEA methodologies.
- PFMEA is particularly relevant to process improvement.
- RCA and FMEA serve different but complementary purposes.
- Corrective actions should address verified causes.
- Effectiveness must be verified after implementation.
- FMEA and Control Plans should be updated when process controls change.
The fundamental RCA sequence is:
Define → Investigate → Identify Causes → Verify → Correct → Verify Effectiveness → Standardize
8. Lesson Learnt / Conclusion
A recurring process problem is rarely solved effectively by treating only its visible symptom.
The Green Belt must develop the discipline to move deeper:
What happened?
→ Why did it happen?
→ What process condition allowed it to happen?
→ What evidence supports the proposed cause?
→ What action will prevent recurrence?
The 5 Whys provides a simple structure for investigating causal chains.
FMEA provides a broader framework for anticipating failure modes, assessing risk, reviewing existing controls, and prioritizing risk-reduction actions.
Neither tool should be used mechanically.
A high-quality analysis requires:
- Clear problem definition
- Process knowledge
- Reliable data
- Logical reasoning
- Evidence-based verification
- Appropriate corrective action
- Follow-up monitoring
The ultimate objective is not merely to explain the failure.
It is to reduce the probability of recurrence and strengthen the process.
Key takeaway:
Do not stop at the symptom. Find the cause, verify it with evidence, act on it, and confirm that the problem stays solved.