Experiment Templates & A/B Test Library

Reusable experiment templates including hypothesis statement, experiment plan, data collection template, success criteria, sample-size guidance, and a simple t-test/checklist for small-sample shopfloor experiments — plus a practical runbook and example.

Quick welcome

Run fast, shopfloor-friendly experiments that produce reliable learning. These templates help you state a clear hypothesis, plan how to run and measure the test, choose reasonable sample sizes, and analyze results without overcomplicating the math.

Why this matters

Too many frontline experiments fail because the hypothesis is fuzzy, measurements are inconsistent, or the sample is too small to detect a meaningful difference. These templates keep experiments small, repeatable, and focused on decisions you can act on.

How to use the templates

  1. Fill the Hypothesis Template to sharpen the question.
  2. Complete the Experiment Plan so everyone knows roles, timing, and how work is assigned.
  3. Use the Data Collection Plan to make measurement reliable and simple.
  4. Set explicit Success Criteria before you start.
  5. Run the experiment, collect data, and follow the Analysis Checklist.

Templates

Hypothesis template (fill in)

If we [change X: clear intervention], then [measurable outcome Y] for [unit, e.g., part, batch, shift] will [direction: increase/decrease] by [expected magnitude or qualitative change] within [time window]. We will know the change worked if [success criteria].

Experiment plan (fields to fill)

  • Title
  • Objective — short statement of what you want to learn
  • Owner — who runs the experiment
  • Intervention(s) — describe A (current) and B (new)
  • Units of measurement — part, batch, minute, cycle
  • Scope — lines/shifts/locations included
  • Start / End dates
  • Assignment method — randomization, alternating shifts, or matched pairs
  • Blocking / stratification — e.g., by operator or machine if important

Data collection plan

  1. Primary metric — exact name, units, and measurement method
  2. Secondary metrics
  3. Sampling frequency — every cycle / every shift / daily
  4. Recorder — who collects and where it’s stored
  5. Data quality checks — how to handle missing or outlier values

Success criteria

Specify an explicit decision rule before running the test. Example: "If average cycle time decreases by at least 10% with no increase in defects over three consecutive days, adopt the change."

Sample-size guidance (shopfloor rules of thumb)

  • For large, obvious effects: 10–20 units per group may show a decisive signal.
  • For moderate effects: 30–50 per group is safer.
  • For small effects or noisy processes: hundreds may be needed — consider redesigning the intervention or reducing noise with blocking.
  • When unsure, prioritize better measurement and blocking rather than blindly increasing sample size.

Simple analysis & t-test checklist

  1. Check data quality (no transcription errors).
  2. Visualize group means with a small table or chart.
  3. Confirm roughly similar variability; if very different, consider a non-parametric check.
  4. Compute group means and difference; calculate a basic t-test if sample >= ~10 per group.
  5. Interpret results against the pre-set success criteria, not only p-values.
  6. Document confidence, practical significance, and any side effects observed.

Worked example (short)

Hypothesis: If we apply the new quick-change checklist, average changeover time will fall by 20% per setup within 2 weeks. Primary metric: minutes per changeover recorded per event. Plan: Run new checklist on alternate shifts for 10 changeovers each (A = current, B = checklist). Success: mean checkouts B at least 20% lower and no defect increase. After collecting 10 events per group, compute means and check difference vs. 20% threshold. If signal is ambiguous, extend run with blocking by operator.

Run checklist (practical)

  1. Brief the team and train the operators on the intervention.
  2. Start data collection and verify first 3 records for correctness.
  3. Stop any changes to other variables during the test window if possible.
  4. Keep a short log of anomalies that could affect interpretation.
  5. Analyze using the Analysis Checklist and decide: adopt, adapt & retest, or abandon.

Common pitfalls

  • Changing too many things at once.
  • Using the wrong unit of analysis (measure per batch when the effect is per part).
  • Ignoring contextual differences between shifts or operators.

Next steps & capability suggestions

This Template Collection is a good candidate for an interactive experiment-planning form that collects plans and stores them for later analysis and organizational memory. An interactive form could guide users through each template field, validate required entries (e.g., primary metric, start date), and save the plan to the experiment history. Pairing that with a simple submission workflow makes experiments easier to repeat and audit.

Image idea: team planning an experiment on shop floor


Discussion

Comments and conversation will live here.