Predictive Maintenance Pilot: Data & Sensor Readiness Checklist
A practical, step-by-step checklist to verify sensors, signals, sampling, labels, sync, data quality, ownership, acceptance criteria, and fallback processes before running a sensor-driven predictive maintenance pilot.
Welcome — why this checklist matters
Before you invest in models or dashboards, confirm the signals, sampling, labels, and operational context are ready. Use this checklist to reduce risk, avoid wasted engineering time, and increase the chance your pilot produces actionable, repeatable value.
How to use this checklist
Walk the list with your reliability engineer, operations lead, controls/automation engineer, and the operator who knows the asset. Mark each item complete, note evidence (files, screenshots, logs) and owners, and don’t start modeling until acceptance criteria are met in the final section.
Asset selection rationale
- Document why this asset: Describe business impact (downtime cost, safety, throughput), frequency of failures, and whether failure modes are diagnosable by sensors.
- Failure-mode clarity: List target failure modes, how they present (symptoms), and typical lead time from first signal to failure.
- Availability of ground truth: Confirm you can label failures from work orders, maintenance logs, or operator reports and estimate number of labeled events historically.
Signal map (what to capture and why)
List candidate sensors and expected value for detecting the target failure modes.
- Vibration (accelerometer): Bearing wear, imbalance, resonance — provide mounting point and orientation.
- Temperature (thermocouple/RTD): Overheating motors, bearings, friction zones — note response time and sensor location.
- Current (clamp-on or CT): Motor load, stalling, electrical faults — include CT ratio and wiring diagram.
- Acoustic/ultrasonic: Leak detection, cavitation, arcing — specify sensor placement and frequency band.
- Pressure/flow: Hydraulic/pneumatic failures or clogging — note sensor tap points and expected operating range.
- Control/PLC signals & events: Mode changes, setpoint changes, alarm triggers — document tags and event semantics.
- Operator/line events: Shift logs, start/stop times, product changes — capture event IDs used to align with sensor data.
Sampling rates & data resolution
- Specify sampling for each signal: e.g., vibration 1–25 kHz (or appropriate frequency band), temperature 1 sample/sec to 1 sample/min depending on dynamics, current 1–10 kHz for inrush/transients or 1–10 Hz for steady load.
- Explain why: Tie sampling choice to expected signal frequency and the shortest event you must detect (capture Nyquist informally: sample at least twice the highest frequency of interest).
- Temporary higher-rate capture: Plan short-duration high-resolution captures for diagnosing intermittent events if continuous high-rate streaming is impractical.
- File formats & compression: State formats (CSV, Parquet, binary), expected file size, and retention plan.
Timestamping & synchronization with events
- Single timebase: Ensure all data sources use the same timezone and NTP-synced clocks or provide conversion mapping.
- Event alignment: Capture production events, maintenance start/end, alarms, and operator notes with stable timestamps or event IDs.
- Latency awareness: Document transmission delays (edge buffering, network lag) and whether timestamps reflect measurement time or ingestion time.
Historical labels & failure data
- Label source: Identify authoritative label sources (CMMS work orders, OEE logs, operator reports) and map label fields to failure modes.
- Label quality: Confirm labels include start and end times of events and a confidence note if labels are uncertain.
- Minimum examples: Aim for at least dozens of labeled failure events for simple models; if you have fewer, plan rule-based detection or anomaly approaches and extend labeling during the pilot.
- Non-failure examples: Collect representative normal operating data across shifts, product variants, and environmental conditions.
Data quality tests
- Completeness: Check % of missing samples per day (<5% target for pilot).
- Constant-value checks: Flag sensors stuck at one value or repeating a pattern (sensor failure or connection issue).
- Out-of-range values: Identify sensor readings outside plausible operating ranges.
- Signal-to-noise & sensitivity: Visualize example windows for known failures and verify the candidate signal shows a detectable change.
- Drift & calibration: Look for slow drift; confirm last calibration dates and plan calibration during pilot if needed.
- Sampling jitter: Inspect inter-sample time variability and correct timestamps if irregular.
Ownership, access, and alerting rules
- Data owner: Assign an owner for each sensor and the dataset (name, role, contact).
- Access control: Confirm who can read and who can write data. Ensure analysts have a read-only copy for modeling.
- Alert routing: Define who receives alerts, how (email, SMS, existing CMMS ticket), and expected response SLAs during the pilot.
- Escalation: Provide escalation path for persistent false alarms or missed detections.
Acceptance criteria & pilot success metrics
Define measurable criteria before modeling begins. Example pilot metric targets:
- Model performance (if supervised): True positive rate (recall) ≥ 70% and false alarm rate ≤ 10% — adjust for asset criticality.
- Anomaly approach: Detection lead time ≥ X hours before failure in ≥ Y% of labeled events (e.g., lead time ≥ 4 hours in 60% of cases).
- Operational impact: Meaningful work order creation from predictions in pilot window and ≥ Z% of predictions produce useful corrective action (operator confirmation).
- Data health: Data completeness ≥ 95% during the pilot, and signal SNR sufficient to separate failure vs. normal windows.
- Business outcome: Reduction in unplanned downtime minutes or avoidance of a critical failure (compare pilot window to baseline period). Set a realistic target tied to business case.
- Human-in-the-loop: Operators confirm alerts are understandable and produce clear next steps (survey or interview feedback score target).
Fallback manual process during pilot
- Fallback defined: Describe the manual monitoring or inspection steps to take if the predictive system is offline or produces an alert you don't trust.
- Document instructions: Provide a simple checklist or standard work for technicians to follow when an alert occurs (what to inspect, who to notify, temporary lockout/protective steps).
- Safety & compliance: Confirm any intervention adheres to LOTO, safety procedures, and regulatory requirements; do not rely on untested automated actions for safety-critical steps.
Quick validation plan (small, fast experiment)
- Collect a short baseline window (several representative days or weeks) that includes at least one known event if possible.
- Run simple signal checks and a baseline threshold detector or spectral feature extractor to verify failures are detectable.
- Compare predicted events against labeled failures and calculate quick metrics (TP, FP, FN) and lead time.
- If results are promising, run a live shadow phase: issue alerts to a review inbox but do not automate actions; collect operator feedback for 2–4 weeks.
Storage, retention & metadata
- Where data lives: Record storage locations, buckets, access URLs, and folder conventions.
- Metadata: Ensure each dataset contains sensor ID, location, units, sampling rate, calibration date, and owner in metadata files.
- Retention policy: State how long high-resolution data will be kept and when it will be downsampled or archived.
Documentation & artifacts to attach
- Wiring diagrams, mounting photos, and sensor datasheets.
- Sample raw data files (1–2 windows from normal and failure conditions).
- Label mapping file (work-order ID → failure mode → start/end timestamps).
- Baseline dashboards or signal plots demonstrating signal change during failure.
Next steps & scale considerations
- If pilot meets acceptance: Produce a scale-up roadmap: standardize sensors & mounts, integrate prediction outputs into CMMS, define operational SOPs, and plan phased roll-out by asset class.
- If pilot fails on data readiness: Prioritize fixes (mounting, sampling, labeling) and consider a second small pilot focused only on data collection and labeling improvements.
Sign-offs
- Data owner: ____________________ Date: ______
- Reliability lead: ______________ Date: ______
- Operations/plant lead: _________ Date: ______
Discussion
Comments and conversation will live here.