Downtime Root-Cause Runbook

A field-ready containment and root-cause runbook for frontline teams and maintenance leaders. Includes a clear immediate-containment checklist, triage guidance, an incident data capture template, a guided 5‑Why facilitation form, evidence checklist, quick corrective-action templates, a short-term workaround vs durable-fix decision matrix, CMMS entry standards, and a 30/60/90 follow-up tracker to ensure durable outcomes and continuous learning.

Purpose and scope

This runbook helps teams recover production faster while capturing the right data so fixes stop recurring. Use it for any unplanned stoppage, partial stoppage, or repeat fault that affects output, quality, or safety. It is designed for operators, shift leads, maintenance technicians, and supervisors who must contain an incident, gather evidence, and move from temporary containment to durable corrective action.

Quick principles

  • Protect people and product first — then production.
  • Contain fast; investigate methodically.
  • Capture clear evidence while it’s fresh (logs, operator notes, photos, timestamps).
  • Prefer a durable fix when justified; use a documented workaround only as a controlled short-term option.
  • Record everything in the CMMS with standardized fields and tags (see CMMS standards below).

Immediate containment checklist (do this now)

  1. Safety check: Ensure area is safe. Lockout/tagout if required. Stop equipment only as needed to protect people.
  2. Stabilize product: Stop producing defective parts, isolate bad product, and prevent contamination of downstream processes.
  3. Restart plan: If restarting is part of containment, follow Standard Work for safe restart (note operator/tooling setup, settings).
  4. Assign roles: Designate Incident Lead, Operator point person, Maintenance technician, and Recorder.
  5. Short-term workaround: Implement a documented workaround only if it keeps safe production and clearly marks the line/pieces for inspection (see decision matrix).
  6. Notify: Inform shift supervisor, production planner, and quality as required by SOP.

Triage guidance (quick decision flow)

Use this triage logic to prioritize response:

  • If safety at risk → Stop and follow safety protocols, escalate immediately.
  • If product safety/quality at risk → Stop production and isolate product; notify Quality.
  • If production loss > threshold (e.g., 30 minutes or X units) or repeated fault → Treat as high priority and escalate to Maintenance Supervisor.
  • If fault is single, transient, and recovery verified → Document and monitor for recurrence.

Incident data to capture (record these fields)

Capture structured data quickly. Either record in the CMMS or on a printed incident card to upload later.

  • Incident ID (generate or CMMS ticket number)
  • Date and exact start time / stop time
  • Machine/asset ID, location, line
  • Shift and crew
  • Operator name and Maintenance technician(s)
  • What happened (short plain-language description)
  • Observed symptoms (alarms, noises, smoke, vibration, parts jammed)
  • Immediate actions taken (containment, temporary adjustments)
  • Parts replaced / settings changed / tools used
  • Evidence collected (photo filenames, log excerpts, test readings)
  • Time to recovery and production loss estimate
  • Tags: recurring-fault, safety, quality, root-cause-investigation-required

Evidence checklist

Collect evidence while fresh. This is vital for accurate root cause analysis.

  • Photos: overall machine, close-ups of failure area, serial/ID plates, damaged parts (use timestamps)
  • Operator notes: what they saw, last normal action, last changeover, last part batch
  • Machine logs/alarms: export or screenshot alarm history with timestamps
  • PLC/SCADA snapshots or trend data around the event
  • Maintenance logs: last PM, last repairs, parts replaced
  • Material lot numbers or supplier information if relevant
  • Measurements and test readings (temperatures, pressures, electrical readings)

Guided 5‑Why facilitation form

Use the 5‑Why to move from symptom to cause. Facilitate with the Incident Lead, Operator, and Maintenance present. Write answers as facts where possible.

  1. Why did the machine stop? (answer 1)
  2. Why did that happen? (answer 2)
  3. Why did that happen? (answer 3)
  4. Why did that happen? (answer 4)
  5. Why did that happen? (answer 5)

Stop when you reach a root cause that you can reasonably address with a corrective action (a process, material, design, training, or control). If answers become speculative, gather more evidence or use a deeper RCA method (fishbone, fault tree) before deciding.

Quick corrective-action templates (use as starting text blocks)

Use these templates when making CMMS entries or drafting work orders.

  • Temporary workaround: "Implemented [describe workaround] to allow safe operation until permanent fix. Effective from [time]. Owner: [name]. Monitoring: [what to check and frequency]. Expected review: [date]."
  • Immediate corrective action: "Replaced [part] and adjusted [setting]. Verified function by [test]. Time to recovery: [duration]. Owner: [name]."
  • Permanent corrective action (proposed): "Investigate and implement [design/process/training/control change] to prevent recurrence. Scope: [assets/process]. Estimated effort: [hours/cost]. Owner: [name]. Target completion: [date]."

Workaround vs durable-fix decision matrix (criteria)

Use these questions to decide whether to accept a short-term workaround or proceed immediately to a durable fix.

  • Does the workaround preserve safety and product quality? If no → durable fix required before restart.
  • Is expected downtime for a durable fix longer than acceptable (production loss > threshold) and can the workaround be safely controlled? If yes → use workaround with strict controls and short review window.
  • Is this a recurring fault (tagged recurring-fault)? If yes → prioritize durable fix.
  • Does the durable fix require engineering or capital spend that must be prioritized? If yes → document in capital plan but implement interim controls and defined review dates.

CMMS entry standards (minimum fields and tags)

Standardized CMMS entries let us trend and find repeaters. At minimum include:

  • Title: short plain-language summary
  • Incident ID / Ticket number
  • Asset ID
  • Failure category (mechanical, electrical, process, human/operator, material)
  • Severity (minor / significant / critical)
  • Downtime minutes
  • Root-cause category (after RCA)
  • Corrective action type (temporary / immediate / permanent)
  • Owner and target completion date
  • Attachments: photos, logs, screenshots
  • Tags: recurring-fault, safety, quality, customer-impact

30/60/90 day follow-up tracker (ensure durable outcome)

Schedule short reviews to confirm whether corrective actions held and to capture learning.

  • 30 days: Verify immediate corrective actions and workaround monitoring. Confirm no recurrence. Owner: [name]. Result: [pass/fail].
  • 60 days: Review progress on permanent corrective action. If delayed, document reason and risk mitigation. Owner: [name].
  • 90 days: Confirm permanent fix implemented and trending is stable. Update CMMS root-cause and close ticket. Owner: [name].

Escalation and meeting cadence

For high-priority or recurring incidents:

  • Immediate escalation to Maintenance Supervisor and Production Manager for critical safety or production-impact incidents.
  • Weekly recurring‑fault review: maintenance and production review top 5 repeaters for trend signals and resource planning.
  • Monthly reliability review: prioritize durable fixes into the maintenance backlog or capital plan.

Common pitfalls and tips

  • Don’t skip evidence collection — guesses waste time and cause repeat failures.
  • Document temporary fixes clearly with owner and expiration date.
  • Use standardized failure categories to enable trending across shifts and lines.
  • Engage operators early; they often know the sequence that leads to failure.

When to escalate to deeper RCA

Use a deeper method (fishbone, fault-tree, FMEA, or external engineering analysis) when:

  • Failure recurs after corrective action
  • Multiple assets show similar faults
  • Potential safety or regulatory impact exists
  • Cost of downtime or scrap is high enough to justify formal study

How to use this runbook in the field

Keep a printed laminated card at the line and a digital copy in the operator/maintenance tablet. Fill the incident fields immediately and attach photos in the CMMS before closing the ticket. Use the guided 5‑Why with the team within 48 hours and schedule the 30/60/90 follow-ups when assigning owners.

Image search phrase

machine downtime investigation


Discussion

Comments and conversation will live here.