OEE Loss Diagnosis Quick Guide (Availability, Performance, Quality)
A practical, step-by-step diagnosis flow, concrete data checks, and operator-centered countermeasures to find, contain, and reduce the largest OEE losses in availability (downtime), performance (speed), and quality (scrap). Includes short experiments, ownership suggestions, and 30/90-day actions teams can run and measure.
Purpose and quick claim
This quick guide helps frontline teams rapidly find the single biggest OEE losses, verify the measurement, contain urgent harm, and run short experiments that produce verifiable reductions. It focuses on availability (downtime), performance (speed), and quality (scrap) while keeping operators and timestamped evidence central to diagnosis.
How to use this guide
Use this during a focused OEE huddle or short improvement sprint. Follow the diagnostic flow, capture simple time‑stamped evidence, pick a tractable loss to attack this shift, run a short experiment, and measure results. Prioritize containment first, root cause second, and organizational learning third.
Quick flow (one-line)
- Verify measurement and data sources.
- Triage the top loss type: downtime, speed, or quality.
- Run the focused diagnosis for that loss type (use timestamped events).
- Implement immediate containment and assign an owner.
- Design a short experiment (shift or day) with clear acceptance criteria.
- Measure, learn, and scale a 30/90-day countermeasure plan.
Step 0 — Verify measurement (do this first)
Before diagnosing, confirm your OEE numbers are trustworthy. Bad data leads to wasted effort.
- Confirm time base: Is availability measured against scheduled production time, planned production time, or clock time? Use a consistent definition.
- Check event logging: Are downtime events timestamped automatically (MES/PLC) or manually (paper/logbook)? If manual, verify a sample of events by observing one shift.
- Spot-check calculations: Recalculate OEE for one shift using raw minutes, pieces, and standard cycle times. Match to the reported value.
- Look for optimistic reporting: Are micro-stops hidden inside aggregated categories? Are scrap and rework recorded consistently?
Step 1 — Triage the top loss type
Use a Pareto of time (minutes) lost or lost throughput to pick the largest tractable loss. If you lack data, run a one-shift time-distribution observation: record every downtime, speed loss, and quality reject with a start/end timestamp.
Step 2 — Focused diagnosis by loss type
Downtime (Availability)
Goal: turn long, repeated, or high-frequency stops into temporary containment and then a corrective experiment.
- Collect timestamped events for 3–5 recent occurrences (start, stop, duration, operator note).
- Run a rapid 5‑Whys for each event with the team at the machine. Use the timestamps to anchor causes (e.g., failure after 12 minutes of operation).
- Classify causes: mechanical failure, tooling, setup/changeover, material, utilities, or human/skill issue.
- Containment examples: remove affected parts from line, prepared backup tooling, temporary SOP reminder, call maintenance for interim fix.
Simple data to capture: event id, start time, end time, duration, operator, immediate containment action, suspected cause.
Performance (Speed)
Goal: identify where actual cycle time consistently exceeds standard cycle time or where micro-stops reduce net speed.
- Measure cycle time over a run of 20–50 units (timestamp each part start or use OEE/PLC time series).
- Look for patterns: slowdowns at certain operators, batches, shifts, material lots, or after changeovers.
- Check changeover procedures and takt alignment; run a timed SMED-style trial if changeover appears to be the issue.
- Quick countermeasures: operator coaching, simple poka-yoke for feeding, visual cycle time timers, minor tooling adjustments.
Quality (Scrap / Defects)
Goal: find where defects originate and add containment to prevent escapes.
- Capture defect counts with timestamps and link them to material lot, operator, shift, and machine settings.
- Identify key CTQs (critical-to-quality dimensions) and check the inspection points nearest the source—don’t only inspect at the end of the line.
- Containment examples: segregate suspect lots, increase inspection frequency for one shift, tag and quarantine nonconforming parts, deploy a simple gage or visual check at the point of production.
Quick countermeasure examples with ownership and 30/90-day actions
Downtime — Example: Repeated bearing failure
- Immediate owner: Line supervisor
- Immediate containment (this shift): Keep spare machine ready; instruct operator on temporary re-lubrication every 2 hours and record timestamps.
- 30-day: Root-cause investigation with maintenance (replace bearing type, confirm lubrication spec, install simple failure alert).
- 90-day: Update preventive maintenance schedule, add sensor or simple timer alert, and train operators on lubrication checks.
Performance — Example: Slow cycles after changeovers
- Immediate owner: Process lead
- Immediate containment: Standardize and time the changeover steps, assign one person to prepare next tooling in advance, record changeover time.
- 30-day: Run a SMED mini‑project to separate internal/external tasks and trial quick-change fixtures.
- 90-day: Adopt most effective SMED changes, update standard work, and include in training.
Quality — Example: Intermittent defective joint
- Immediate owner: Quality technician
- Immediate containment: Add a touch-point check right after the joint; quarantine suspect lot.
- 30-day: Test material batch, re-check process parameter settings, run root-cause with cross-functional team.
- 90-day: Implement process control (fixture changes or sensor), update work instructions, and add gage at source.
Designing a short experiment
Keep experiments small, time-bound, and measurable.
- Hypothesis: If we do X, downtime Y will fall by Z minutes per shift.
- Timebox: One shift to one week depending on frequency of events.
- Metrics: Minutes downtime per shift, parts/hour, reject rate; capture baseline for the same shift/operation.
- Acceptance criteria: A clear percentage or absolute improvement (e.g., >30% reduction in downtime minutes across three shifts) and no adverse effects on quality or safety.
- Owner & experiment log: Who runs it, start/end times, observations, and timestamped data files or photos.
Simple 5‑Whys template (use at machine, with timestamps)
- Why did the machine stop? — e.g., bearing failed (time: 09:12)
- Why did the bearing fail? — e.g., insufficient lubrication (09:12–09:14)
- Why was lubrication insufficient? — e.g., lubrication interval missed during last shift change
- Why was the interval missed? — e.g., no visible reminder or documented handover step
- Why is there no reminder? — e.g., changeover checklist missing that step
Turn the last line into an action (add checklist item) and run a one-shift test.
Operator engagement and documentation
Bring operators into every step: data collection, quick containment, hypothesis design, and acceptance criteria. Use simple logs or photos. If possible, capture one video or a few photos during the event to preserve context for the team review.
Common pitfalls to avoid
- Fixing the symptom without timestamped evidence—measure before/after.
- Chasing low-impact metrics—prioritize by time lost or throughput gained.
- Making changes without owner or acceptance criteria—each experiment needs a single accountable owner.
- Relying only on anecdote—combine operator input with a quick sampled dataset.
One-shift checklist (use this on the shop floor)
- Verify OEE numbers for this shift by recalculating one hour of production.
- Record top 3 loss events with timestamps (downtime, slow run, defects).
- Pick one highest-impact loss to contain this shift and assign an owner.
- Run a rapid 5‑Whys at the machine; document with timestamps.
- Define a timeboxed experiment (owner, metric, baseline, acceptance) and run it.
- Capture results and plan the 30/90-day countermeasure if successful.
Next steps and scaling
If the short experiments produce reliable improvements, move to formalize countermeasures: update standard work, preventive maintenance, training, and add simple alerts or checks. For larger or recurring mechanical issues consider sensor-based monitoring or MES integration to ensure automatic timestamping and alerts.
Where this guide can be enhanced by platform capabilities
Converting this guide into an interactive diagnosis workbook (timestamped event form, experiment planner, and countermeasure tracker) would make it easier to capture and reuse shift-level evidence and to roll up results for dashboards and trend analysis.
End of guide — use this as a practical, operator-centered starting point. Measure everything you can and let short experiments prove which countermeasures actually deliver lasting OEE gains.
Discussion
Comments and conversation will live here.