Experimentation & Learning Sprint Template

A practical, fillable experiment template that guides teams from hypothesis through design, data plan, decision rules, and a structured post-mortem. Includes field-level guidance, example decision rules, and a short rubric for assessing learning quality.

Experimentation & Learning Sprint Template

This template helps teams run short, hypothesis-driven experiments that generate reliable learning. Use it to register an experiment, make explicit the metric and decision rules, capture findings, and preserve organizational learning. Keep entries concise but concrete — this is meant to accelerate clear decisions and fewer ambiguous results.

How to use this template

Fill the fields before the experiment begins. Be specific about measurements, sample sizes or duration, and the decision rule that will determine next steps. After the experiment ends, record findings, analysis notes, and a short post-mortem. Save this record for team memory and reuse.

Template fields (with guidance and examples)

  1. Experiment title

    Short descriptive name (e.g., “Checkout flow A/B — reduce friction on shipping selection”).

  2. Owner & collaborators

    Who is responsible for this experiment and who will contribute to design, data, and implementation.

  3. Purpose / Problem

    A 1–2 sentence statement of the problem you hope to address and why it matters.

  4. Hypothesis & expected effect size

    State a falsifiable hypothesis and the minimum practical effect you expect (e.g., “If we auto-select the recommended shipping, conversion will increase by ≥ 2 percentage points”).

  5. Primary metric (definition)

    Exactly how the metric is calculated (numerator, denominator, filters). Example: “Checkout conversion = number of successful orders / number of sessions entering checkout, excluding gift-cards-only orders.”

  6. Baseline value

    Current value of the primary metric (with time window and sample). Example: “Baseline conversion = 8.4% (last 30 days).”

  7. Minimum detectable change (MDC) or minimum worthwhile effect

    What change you consider practically meaningful. If you don’t have a formal MDC, state a sensible threshold for decision-making (e.g., +1.5 percentage points). Note: if precision matters, compute sample size or MDC using a basic power calculation or a sample-size calculator.

  8. Design summary

    Describe sample (population, targeting), randomization method, variations (control, variant names), allocation ratios, and any blocking or stratification. Include duration or sample-size target. Example: “50/50 A/B, all desktop users in US, expected 14 days to reach N=20k sessions per arm.”

  9. Data collection & instrumentation checklist

    List events, tags, and metrics to collect. Confirm instrumentation is deployed and validated before starting. Note any known measurement limitations.

  10. Analysis plan

    Describe the primary statistical approach (e.g., difference in proportions with 95% CI, non-parametric test), handling of multiple comparisons, and whether interim peeking is allowed. Keep it simple and pre-specified to avoid bias.

  11. Decision rule (pre-specified)

    Explicit criteria that map observed results to actions (Go / No-Go / Pivot / More Data). Example: “If variant increases primary metric by ≥ 1.5 pp and 95% CI excludes 0, rollout to 100%.”

  12. Risk, safety & compliance checks

    Identify user safety, legal, privacy, or operational risks. State mitigation steps and approval status.

  13. Start & planned end dates

    When the experiment starts and the planned stop condition (date OR sample size OR duration OR decision rule triggered).

  14. Findings (post-run)

    Summarize results against the primary metric (effect size with uncertainty), secondary observations, any anomalies, and data-quality notes.

  15. Decision taken

    Record the action chosen (Go / No-Go / Pivot / More Data) and who approved it.

  16. Next steps & rollout plan

    Concrete actions, owners, timeline, and monitoring plan for rollout or follow-up experiments.

  17. Key lessons & knowledge capture

    Short list of what was learned, surprising insights, and recommended hypotheses for follow-ups. Keep this practical for teammates who were not involved.

  18. Tags / domain / product area

    Useful for indexing and later reuse (e.g., payments, onboarding, labelling).

Post-mortem rubric (quick quality check)

Use these dimensions to rate whether the experiment produced actionable learning:

  • Clarity (Was the hypothesis & metric clear?) — Yes / Partial / No
  • Design (Was allocation and instrumentation appropriate?) — Yes / Partial / No
  • Data quality (Were measurements trustworthy?) — Yes / Partial / No
  • Decision (Was the decision rule applied objectively?) — Yes / Partial / No
  • Learning (Did we learn something usable?) — High / Moderate / Low

Flag any “Partial” or “No” answers and capture corrective actions before reuse.

Example decision rules

  • Go: observed lift ≥ 2.0 pp and 95% CI excludes 0; product team greenlights rollout.
  • No-Go: observed lift ≤ 0 or negative impact on secondary safety metric; experiment aborted.
  • More Data: observed lift in (0, 2.0) with CI crossing 0; extend duration or increase sample size.
  • Pivot: primary metric unchanged but a strong signal on an unexpected secondary metric worth exploring.

Notes & tips

  • Pre-specify as much as practical to reduce bias. Avoid changing primary metric mid-run unless there is a documented measurement failure.
  • Use sensible MDC values tied to business impact, not purely statistical significance.
  • When in doubt about sample size, run a short pilot to validate instrumentation and variance estimates, then compute a sample-size target.
  • Capture lessons in a searchable registry so future teams learn what worked and what didn't.

Where this fits in the "Becoming Your Best" domain

This template addresses the hunger: “Run fast, safe experiments that generate reliable learning and reduce risk.” It reduces the mal-hunger of wasted effort and ambiguous results by making decisions and measurements explicit, repeatable, and discoverable.

Suggestion: Consider converting this template into an interactive experiment registry form so teams can save structured experiment records, run audits, and build dashboards. See Capability Enhancement notes for options.


Discussion

Comments and conversation will live here.