Evidence-Quality Rubric — quick checklist

Before you commit to scaling or stopping an experiment, run this short rubric. Score each dimension 1 (weak) to 5 (strong). If the average is below 3, treat the finding as provisional and plan mitigation (more data, replication, improved protocol).

Rubric dimensions

  1. Design & controls — Were controls appropriate? Was randomization or counterbalancing used where relevant?
  2. Sample size & precision — Is the sample size sufficient to detect the expected effect with reasonable precision?
  3. Measurement validity — Are the instruments or assays validated and reliable for the outcome?
  4. Pre-registration & analysis plan — Was the analysis plan defined ahead of time to reduce bias?
  5. Reproducibility / replication — Has the observation been independently replicated or internally reproduced?

Scoring & interpretation

Score each dimension 1–5 and compute an average:

  • Average < 3: Weak evidence — do not scale. Plan replication or improved measurement.
  • Average 3–4: Mixed — proceed with caution and clear replication steps before investment.
  • Average > 4: Strong evidence — consider scaling or embedding into the roadmap with monitored KPIs.

Example (short)

Metric: Process yield improved by 5% (n=12). Scores: Design 3, Sample size 2, Measurement 4, Pre-registration 1, Replication 2. Average = 2.4 → treat as provisional and replicate with larger n and documented analysis plan.

How to use it in the huddle

Ask the metric owner to provide a 30–60 second evidence summary and the rubric scores (or use this checklist to score together). Use the result to drive the binary decision (continue/stop/pivot/scale) and specify the next step (replicate, change measurement, escalate).


Discussion

Comments and conversation will live here.