Data Quality & Governance Audit Toolkit

An interactive, structured audit that helps research teams assess data completeness, lineage, schema validation, access and retention controls, stewardship, and remediation priorities — with saved responses to track findings and follow-up.

Interactive Tool

Data Quality & Governance Audit

Use this guided audit to rapidly assess the health and governance of critical research datasets. Capture findings, assign stewards, and prioritize remediation by scoring impact and effort. The form saves your responses so teams can track progress and reproduce decisions.

Before you begin, gather a list of critical datasets, any available metadata or lineage diagrams, and an inventory of current access and retention policies.

Describe the boundaries of this audit: teams, databases, file systems, folders, experiments, time range. List dataset identifiers if available.
One per line. Include dataset name, owner, location (storage URI), and short purpose (e.g., 'raw sequencing reads for Project X').
Assess the risk that these datasets could threaten analysis or reproducibility if issues persist.
Lineage includes source systems, transformations, derived tables, and consumers.
Choose the best match for current lineage documentation.
If you have automated completeness checks, enter the percent of records/files passing validation (0-100). If unknown, leave blank.
Indicates whether structural and type checks are enforced when data is ingested or transformed.
Consider who has read/write/admin permissions and whether principle of least privilege is enforced.
Summarize misconfigurations, excessive privileges, public exposures, or missing approvals.
Does current retention practice match policy and regulatory requirements?
Name and contact for the person accountable for dataset quality and governance.
Has the steward been informed and agreed to take ownership?
Rate 1 (low) to 5 (severe). Consider reproducibility, patient safety, regulatory risk, or business impact.
Rate 1 (low) to 5 (very high). Consider people, time, and technical complexity.
High priority typically means high impact and low-to-medium effort.
Be specific: e.g., 'Add schema enforcement in ingestion pipeline', 'Re-run validation for dataset X', 'Remove public read permission', 'Document lineage for transformation Y'. Include estimated owner and timeline.
Enter a date or quarter. This field is for planning only.
Name and contact for the person who will track remediation progress.
Attach or link to lineage diagrams, validation logs, policy documents, issue tracker IDs, or screenshots.
Your overall assessment of dataset governance health: 1 (poor) to 5 (excellent).
Anything else the team should know, context, constraints, or follow-up suggestions.
You can explore this tool now. Sign in or create an account to save your responses and return to them later.
Make this tool part of your work

Save a personal copy, bring it to your team, or tailor the questions and workflow to fit what you are hungry to improve.

Member customization and team collaboration are coming soon.

Discussion

Comments and conversation will live here.