Reliability RCA Package: FMEA → RCM → Corrective Actions

A compact, practical playbook and toolkit that turns repeat failures into durable fixes. Includes a failure‑mode prioritization worksheet (FMEA-lite), a guided RCA workflow, a solution evaluation matrix, and a post‑implementation verification checklist with acceptance criteria and examples.

Purpose

This toolkit helps maintenance, engineering, and operations teams systematically prioritize equipment failure modes, perform focused root cause analysis (RCA), choose robust corrective actions, and verify that actions prevent recurrence. It is designed for teams that work with imperfect CMMS data and limited resources—practical, repeatable, and adaptable to small shops or large plants.

How to use this toolkit

  1. Gather a short list of recurring failures (top 5–10) from logs, operator reports, and tribal knowledge.
  2. Use the Failure‑Mode Prioritization Worksheet to score and rank items for focused effort.
  3. Run a structured RCA for the top priorities using the step‑by‑step RCA guide.
  4. Evaluate candidate solutions with the Solution Evaluation Matrix to balance effectiveness, cost, risk, and feasibility.
  5. Implement the chosen action, document required work packages or design changes, and assign ownership.
  6. Use the Post‑Implementation Verification Checklist to confirm the problem is resolved and improvements are sustained.

Failure‑Mode Prioritization Worksheet (FMEA‑lite)

Use a short FMEA-style table to prioritize. Keep it practical: limit to essential columns so teams actually score items.

Item / Asset Failure Mode Frequency (1–5) Impact (1–5) Detectability (1–5) Priority Score (F × I ÷ D) Notes / Sources
Example: Press #2 Hydraulic leak → unplanned stop 4 4 2 8 Operator logs, last 3 months

Scoring guidance: Frequency = how often it occurs; Impact = production, safety, quality, cost; Detectability = how likely the failure will be detected before it causes a stoppage. Use consistent scoring scales across the site.

Step‑by‑Step RCA Guide

Run a focused RCA workshop (30–90 minutes) with people who know the process. Use both evidence (data, logs, photos) and structured inquiry.

  1. Define the problem clearly. Write a one‑sentence problem statement with WHEN, WHERE, WHAT, and IMPACT.
  2. Containment (immediate action). What temporary measures keep production safe/reliable while you investigate? Assign containment owner and deadline.
  3. Collect evidence. CMMS history, inspection records, operator notes, photos, wear measurements, environmental readings, and failed parts retained for analysis.
  4. Use dual analysis methods. Start with 5 Whys for a quick causal chain, then validate with a Fishbone (Ishikawa) to explore mechanical, electrical, material, human, and process causes.
  5. Identify root cause(s). Prefer proximate actionable causes tied to system, design, or process rather than blaming individuals.
  6. Create corrective action options. For each root cause, list options that remove the cause, reduce the likelihood, or improve detection.

RCA checklist prompts: Who was last to work on the asset? What changed before the failure? Were warning signs missed? Were PM tasks followed? Was the right spare used?

Solution Evaluation Matrix

Score candidate solutions on criteria so decisions are explicit and auditable.

Solution Effectiveness (1–5) Cost (1–5, lower is better) Implementation Time (days) Operational Risk (1–5) Maintainability / Training (1–5) Net Score / Recommendation
Replace seal with upgraded design 5 3 2 2 4 High — recommended

Notes: Weight the criteria to reflect site priorities (safety and downtime reduction usually weigh heaviest). Document who approved the chosen solution and the expected benefit (e.g., MTBF increase, reduced stops per month).

Post‑Implementation Verification Checklist

  • Implementation completed and documented (date, owner, parts, procedures).
  • Workpack or design change record stored in CMMS/engineering repository.
  • Monitoring plan defined (what to measure, frequency, owner, threshold).
  • Short‑term verification: no repeat failure within agreed observation window (e.g., 30/60/90 days depending on cycle).
  • Long‑term check: trending of key indicators (stops/month, MTBF) for 6 months.
  • Training delivered to affected operators/maintenance staff; competence confirmed.
  • Lessons learned recorded and shared (what worked, what didn’t, suggested follow‑ups).

Accept criteria: Define what success looks like (e.g., 50% reduction in stops over 3 months, or elimination of specific failure mode for 90 days). Do not mark closed until acceptance criteria are met and verified.

Short Example (illustrative)

Problem: Conveyor motor overheats and trips weekly, causing 2 hours of downtime.

  1. Prioritization: High frequency (5) × high impact (5) ÷ detectability (2) → high priority.
  2. RCA: 5 Whys revealed cause: Improper bearing lubrication frequency & a low quality grease spec.
  3. Solutions evaluated: revise lubrication interval + better grease (low cost, quick), replace bearing with sealed unit (higher cost, longer lead time).
  4. Selected action: immediate change to grease spec and update PM; plan sealed bearing replacement as capital work if problem persists.
  5. Verification: no trip in 60 days; MTBF improved; update PM and parts list in CMMS.

Adaptation & Practical Tips

  • Keep artifacts short and linked: store completed worksheets in CMMS or the site knowledge base for reuse.
  • If CMMS data is poor, combine operator logs, supervisor notes, and production records to estimate frequency.
  • Use simple templates and time‑boxed workshops so RCA becomes part of normal work rather than a rare audit event.
  • Assign clear owners and due dates for containment, corrective action, and verification.

Recommended next steps

  1. Download or copy these templates into your team space and run the process on one high‑impact failure this week.
  2. Consider turning the worksheets into interactive forms so teams can submit RCA outputs, track verification, and tie to CMMS work orders.
  3. Package the top solved RCAs into a local knowledge collection for onboarding and continuous improvement huddles.

End of toolkit.


Discussion

Comments and conversation will live here.