Evidence-Quality Rubric — quick checklist
Before you commit to scaling or stopping an experiment, run this short rubric. Score each dimension 1 (weak) to 5 (strong). If the average is below 3, treat the finding as provisional and plan mitigation (more data, replication, improved protocol).
Rubric dimensions
- Design & controls — Were controls appropriate? Was randomization or counterbalancing used where relevant?
- Sample size & precision — Is the sample size sufficient to detect the expected effect with reasonable precision?
- Measurement validity — Are the instruments or assays validated and reliable for the outcome?
- Pre-registration & analysis plan — Was the analysis plan defined ahead of time to reduce bias?
- Reproducibility / replication — Has the observation been independently replicated or internally reproduced?
Scoring & interpretation
Score each dimension 1–5 and compute an average:
- Average < 3: Weak evidence — do not scale. Plan replication or improved measurement.
- Average 3–4: Mixed — proceed with caution and clear replication steps before investment.
- Average > 4: Strong evidence — consider scaling or embedding into the roadmap with monitored KPIs.
Example (short)
Metric: Process yield improved by 5% (n=12). Scores: Design 3, Sample size 2, Measurement 4, Pre-registration 1, Replication 2. Average = 2.4 → treat as provisional and replicate with larger n and documented analysis plan.
How to use it in the huddle
Ask the metric owner to provide a 30–60 second evidence summary and the rubric scores (or use this checklist to score together). Use the result to drive the binary decision (continue/stop/pivot/scale) and specify the next step (replicate, change measurement, escalate).
Discussion
Comments and conversation will live here.