Crisis Response Runbook and Plant Recovery Checklist
A concise, actionable plant-level runbook for major disruptions with clear activation criteria, leadership roles, immediate safety and containment checklists, customer and supplier communication templates, step-by-step recovery and restart sequences, quality verification steps, and a post-mortem timeline.
Overview
This playbook is a plant-level runbook for responding to major disruptions (plant outages, utility failures, supplier failure, severe equipment damage, natural disasters, or other incidents that threaten production or safe operation). Its purpose is to help teams act quickly and consistently: secure safety, communicate effectively, contain damage, re-establish critical systems, resume production in a controlled way, and learn so the event doesn't repeat.
How to use this runbook
- Keep one tailored copy per site with site-specific contacts, suppliers, threshold values, and equipment lists.
- Print a one-page Quick Reference for shop-floor supervisors and post it in the control room.
- Run tabletop exercises at least annually and after major process, personnel, or supplier changes.
Activation criteria
Activate the Crisis Response Runbook when any of the following are true:
- Expected or unexpected downtime exceeds X hours for a critical line (site-specific threshold).
- Loss of a major supplier or raw material that will stop production within Y hours/days.
- Major safety incident, fire, explosion, structural damage, or confirmed environmental release.
- Widespread IT/MES/SCADA failure that prevents controlled production or tracking.
- Severe weather or external event causing evacuation, shelter-in-place, or utility interruption.
Crisis leadership roles and responsibilities
Define named alternates for each role. Use clear single-point ownership during the event.
- Incident Commander (IC) – Overall decision authority for the site. Declares activation, prioritizes safety vs. production trade-offs, authorizes external communications.
- Safety Lead – Ensures personnel safety, triage, evacuation, permits re-entry, coordinates with EMS/authorities.
- Operations Recovery Lead – Leads technical recovery: sequences restarts, coordinates maintenance, validates equipment readiness.
- Maintenance Lead – Directs diagnostics, spare parts, and vendor technicians. Tracks repair progress and parts ETA.
- Quality Lead – Defines hold/release criteria, leads sampling plan and testing, approves product disposition.
- Communications Lead – Manages internal communications, customer notifications, and media statements. Keeps stakeholders informed with agreed cadence.
- Supply Chain Lead – Executes supplier contingencies, inventory reallocations, and alternate sourcing.
- IT/Controls Lead – Validates MES/SCADA/PLC integrity, backups, and data preservation.
Immediate safety & containment checklist (first 30–60 minutes)
- Confirm human safety. Triage injured and call emergency services if needed.
- If required, initiate evacuation / lockout and account for all personnel.
- Stop affected processes safely using standard shutdown procedures (stop sequences, isolate energy sources, LOTO where needed).
- Secure hazardous materials and isolate utility feeds (gas, steam, electricity) where appropriate.
- Preserve evidence for later investigation (take photos, secure logs) while maintaining safety.
- Assign Safety Lead to coordinate with local emergency responders and regulatory reporting.
Immediate communications
Communicate early, often, and honestly. Use short, tested templates and log every message.
Initial internal notification (example)
Subject: Plant incident — [Brief description] — [Date/Time]
Body: We have experienced [brief description]. Safety status: [All safe / injuries reported]. Incident Commander: [Name]. Current action: [Evacuation / Containment / Shutdown]. We will provide an update at [time]. Do not post to social media. [Contact name/phone].
Customer initial template (example)
Dear [Customer Name],
We experienced an operational disruption at our [Plant Name] on [Date/Time]. We are preserving safety and assessing production impact. At this time we estimate [initial impact statement]. We will provide a status update by [time]. Your sales/ops contact: [Name, phone, email].
Media/External template (only Communications Lead)
We acknowledge an incident at our [Plant]. Our priority is the safety of everyone on site. We are working with emergency services and will share verified updates as they become available. [Non-speculative factual detail].
Supplier contingency triggers and actions
- Trigger: Critical supplier fails to deliver and on-hand cover < threshold. Action: Supply Chain Lead executes alternate supplier list, places expedited orders, and initiates emergency sourcing playbook.
- Trigger: Long lead time part required for repair. Action: Evaluate cross-plant spares, reallocate inventory, and coordinate expedited vendor shipment or onsite vendor repair.
Recovery checklist: systems and validation
Follow a controlled, documented restart sequence. Never bypass safety interlocks. Record each step and sign off.
- Confirm safety clearance from Safety Lead and IC to begin recovery.
- Restore utilities and building systems (electric, compressed air, water, HVAC) and validate stable supply.
- Restore IT/MES/SCADA backups; validate control system communications and alarms with IT/Controls Lead.
- Bring critical support systems online (pumps, fans, compressors) before process equipment.
- Perform mechanical and electrical inspections on critical equipment (bearing temps, lubrication, alignment, torque checks).
- Start up a single pilot line or cell using controlled settings; monitor key process variables and quality metrics.
- Quality Lead to perform sampling and testing per predefined hold-and-test plan before releasing product.
- Gradually ramp production following checklist thresholds and document deviations.
Quality verification & product disposition
- Use pre-defined sampling plans and test methods. Hold affected lots until cleared by Quality Lead.
- If product integrity is uncertain, quarantine and label for investigation; do not release to customers without approval.
- Record tags, batch numbers, timestamps, and any altered process parameters for traceability.
Post-mortem and corrective actions
- Within 48 hours: initial debrief with leadership to capture timeline and immediate corrective actions.
- Within 7–14 days: formal root cause analysis (5-Why, Fishbone, or equivalent) with cross-functional team, define corrective and preventive actions (CAPAs) with owners and due dates.
- Track CAPAs publicly in the site action board; validate effectiveness and close only after verification of sustained change.
Tabletop exercises & maintenance of the runbook
Schedule tabletop exercises at least annually and after major changes. Exercises should:
- Use realistic scenarios that test communications, supplier contingencies, and restart procedures.
- Include participation from all crisis roles and key external contacts (vendors, customers, emergency responders when practical).
- Capture lessons and update the runbook within 30 days of the exercise.
Quick reference: essential site information (to keep updated)
- Incident Commander & alternates (name, role, phone)
- Local emergency contacts (fire, police, EMS)
- Key vendors and their emergency tech contact info
- Critical spare parts list and storage location
- Site thresholds for activation (downtime hours, inventory days of cover)
- Location of main E-stops, utility shutoffs, and evacuation assembly points
Versioning & ownership
Runbook owner: [role/title]. Update cadence: review annually or after any incident, supplier change, or major process change. Record revision history at the top of the site copy.
Discussion
Comments and conversation will live here.