Data Instrumentation & Quality Checklist for Discovery

Practical interactive checklist to confirm experiments and analytics have the necessary instrumentation, telemetry, sample checks, retention, and access controls before running tests or publishing datasets for discovery.

Interactive Tool

Data Instrumentation & Quality Checklist for Discovery

Use this checklist before running experiments, sharing datasets for discovery, or publishing analytics. It helps teams confirm that instrumentation, identifiers, sampling, logging, retention, and access controls are in place so experiments are reliable, datasets are reusable, and privacy/compliance risks are managed.

Complete the checklist for the dataset/event stream you plan to use. Save the record so teams can track readiness and follow up on any gaps.

A short name or identifier for the dataset, telemetry stream, or experiment instrumentation being evaluated. Use the canonical name used in your catalog.
Date of this checklist (YYYY-MM-DD).
Person completing this checklist and their role (e.g., analyst, SRE, product manager).
Are events, actions, and attributes consistently named and documented so different teams interpret data the same way?
Does each entity (user, device, session, order, etc.) have a stable unique identifier across sources and time?
Are timestamps available, in a documented format, and normalized across producers (including timezone handling)?
Have you defined checks to ensure samples used for experiments reflect the intended population (sampling bias, coverage, and population shifts)?
Are payload schemas, max sizes, expected fields, and error-handling documented and enforced?
Select the automated checks that are configured for this dataset.
Are retention periods, archival locations, and deletion procedures defined and documented for this dataset?
Are roles, access levels, and approval flows defined so only authorized users can query or export data?
Have PII/PHI been identified and controls (masking, pseudonymization, consent flags) been applied or documented?
Do you know which teams, reports, or models will consume this dataset and for what purposes?
Will your platform create alerts when ingestion drops or schema changes occur?
Primary owner and secondary contacts (emails or teams) for follow-up and incident response.
Quick confidence rating for using this dataset in discovery or experiments.
1.0 10.0
Document any known issues, next steps, or owners for items marked Partial or Missing.
You can explore this tool now. Sign in or create an account to save your responses and return to them later.
Make this tool part of your work

Save a personal copy, bring it to your team, or tailor the questions and workflow to fit what you are hungry to improve.

Member customization and team collaboration are coming soon.

Discussion

Comments and conversation will live here.