About this article
As the seventh installment of the “Data Architecture” category in the series “Architecture Crash Course for the Generative-AI Era,” this article explains data governance.
A technology-only data platform rots in 3 years. This article covers the components of governance - data catalog, metadata, lineage, quality management, data stewards, and access control - alongside an introduction roadmap by scale and regulation, and the structure where in the AI era governance functions as a dictionary for AI.
Before you read this
This article uses a good deal of vocabulary from around databases. If that is unfamiliar, reading the primer "Database Basics" first makes it far easier to follow. You can also look anything up in the glossary as you read.
What is data governance in the first place
Data governance is “the mechanism for establishing and continuously enforcing rules on who can use what data and how across the entire company.”
Imagine a condo association. If residents dump trash whenever they want, leave personal belongings in shared areas, and copy keys freely, the whole building falls apart. It’s the management regulations (governance) and the building manager (data steward) that maintain order. Data is the same - without rules for definitions, naming, access permissions, and quality standards, different departments produce different numbers, personal information leaks, and AI learns from bad data. That’s the collapse governance prevents.
Why data governance is needed
Build a data platform and leave it alone, and over time it accumulates “tables nobody knows who made” and “columns with the same name and different definitions” until the platform itself cannot be trusted. Once departments start using data on their own separate definitions, the revenue figure differs by department, and feeding data with unclear definitions to an AI sharply raises the risk of a wrong decision.
The weight of regulation has also changed by an order of magnitude. In May 2023 Meta was fined EUR 1.2 billion under the GDPR for transferring EU citizens’ data to the United States. GDPR fines run up to 4 percent of annual global turnover — this is an era in which “where the data is stored” alone can produce a penalty at the level of company survival.
The main components
To implement data governance, combine the following elements. Any one alone is insufficient - it functions as a trinity of organization, institution, and technology.
| Element | Role |
|---|---|
| Data catalog | Catalog of where what is |
| Metadata management | Each data’s definition, owner, update frequency |
| Lineage | Visualization of data transformations and flow |
| Quality management | Detection of definition violations, missing data, duplicates |
| Access control | Permissions on who can see what |
| Data stewards | Person responsible for each data |
| Policy | Rules for retention period, encryption, disposal |
Catalogue and metadata — getting rid of “where is this data?”
A data catalog is the organization’s catalog of all data, with the mechanism of central searchability for “where, what data, with what definition, managed by whom.” It originated when Google built an internal mechanism to “search data like Google Search,” and today commercial and OSS tools are provided by various vendors.
Without a catalog, analysts ask people every time “where is this data?” and “what is this number?”, dropping data utilization speed to 1/10.
| Pros | Cons |
|---|---|
| Faster data discovery | Heavy initial metadata setup |
| Shorter new-hire onboarding | Information rots if neglected |
| Becomes evidence for audits | Tool fees, operational cost |
| Premise for AI integration | Doesn’t function without a steward regime |
At small scale, dbt docs is enough. Consider dedicated tools from mid-scale onward.
There are several catalogue tools to pick from.
| Tool | When to choose |
|---|---|
| dbt docs | Already on dbt, analytics-model-centric. Lightweight and free |
| DataHub (LinkedIn OSS) | Mid-to-large, want to grow OSS while customizing |
| Amundsen (Lyft OSS) | Lighter and simpler than DataHub. Want to lower the entry bar |
| Apache Atlas | Legacy DWH environment centered on Hadoop/Hive |
| Collibra | Enterprise, want to seriously build a governance regime |
| Alation | Want AI features (natural-language search etc.) |
DataHub is becoming the OSS de facto. Commercial-side has the two giants Collibra (full-feature) and Alation (strong AI), with investment scales for large enterprises.
Metadata management sits alongside it.
The contents of the data catalog are metadata (data about data). For each table/column, record “what it’s for,” “how to use it,” and “who to ask,” so users don’t get lost.
| Metadata type | Content |
|---|---|
| Business metadata | Definition, meaning, business glossary |
| Technical metadata | Schema, types, constraints, indexes |
| Operational metadata | Update frequency, SLA, job history |
| Owner info | Data steward, contact |
| Quality metadata | Test results, anomaly-detection scores |
Metadata needs both auto-collection (catalog tool scans) and manual entry (steward writes), and full automation is impossible.
Lineage makes the flow visible.
Data lineage is the mechanism for tracking where data came from, how it’s transformed, and where it flows. The DAG auto-generated by dbt is one form of lineage, visualizing what columns of which business DB this aggregated result is derived from.
| Use case | Content |
|---|---|
| Impact investigation | ”What breaks if I change this column?” |
| Cause tracking | ”Where did this dashboard’s number go wrong?” |
| Audit response | ”Where is personal data flowing to?” |
| Data retirement | ”Who’s using this table?” |
Without lineage in place, you can’t delete old tables, and mystery tables pile up.
Quality management and stewards — the technical and human halves
Data quality is measured along 6 viewpoints. Auto-testing these and validating during pipeline execution is the modern best practice. dbt’s tests feature and Great Expectations are used.
| Viewpoint | Meaning | Example |
|---|---|---|
| Completeness | No missing data | No NULL in required items |
| Uniqueness | No duplicates | User IDs not duplicated |
| Accuracy | Matches reality | Sales amounts match actuals |
| Consistency | No logical contradictions | end_date > start_date |
| Timeliness | Updates aren’t lagging | Daily data updated daily |
| Referential integrity | Foreign keys exist | User ID exists |
Data without guaranteed quality becomes the worst risk of misleading executive decisions.
The human half is the data steward.
A data steward is a human role taking responsibility for each data set. The organization needs people deciding the things technology alone can’t solve - “what’s this data’s definition?” and “how may it be used?”
| Role | Responsibility scope |
|---|---|
| Business steward | Manage definition, use, business terms |
| Technical steward | Schema, pipeline, quality |
| Data owner | Final approval, scope of disclosure |
| Data custodian | Daily operations, access grants |
Stewards are assigned at least one per data set to clarify the locus of responsibility. The basis of governance is not leaving “tables no one manages” alone.
Access control — down to row and column level
Data containing personal/confidential information is strictly managed for who can access. Mere “table-level permissions” aren’t enough - row-level and column-level access control becomes necessary.
| Control level | Content |
|---|---|
| Table level | Read/write permission per table |
| Column level | Salary column visible only to HR |
| Row level | See only your department’s data |
| Dynamic masking | Mask PII (Personally Identifiable Information) at query time |
| Retention period | Auto-delete after N days |
BigQuery and Snowflake natively support column-level and row-level policies, and SQL-side automatic filtering removes the need for app-side individual implementation.
How to choose — a phased roadmap by scale and regulation
Going “large-scale straight away” stalls operations, so growing it in stages is the realistic path.
| Phase | Organisation size | What you introduce |
|---|---|---|
| 1. Minimum | up to 30 people | dbt docs plus dbt tests |
| 2. Basic | up to 300 people | + named data stewards, naming conventions |
| 3. Mid-size | up to 3,000 people | + DataHub or Amundsen, row- and column-level access control |
| 4. Enterprise | 3,000 and up | + Collibra or Alation, a dedicated governance function |
| 5. Regulated or listed | any size | history retention, audit logs, full lineage |
Regulatory requirements go in at the start. Personal data in the EU falls under GDPR (right to erasure, consent management), healthcare under HIPAA, payments under PCI DSS, listed companies under SOX. In heavily regulated industries governance investment is not optional, and failing to meet it can mean withdrawing from the business. Requirements also rise with how far you take AI: catalogue and consistent naming for BI alone, organised metadata for Text-to-SQL, lineage and quality tests for RAG and agents, and audit logs and explainability for autonomous decisions.
Three scenarios
If you are building solo or at a startup
Starting with dbt docs and dbt tests alone is enough. You need neither a dedicated person nor a dedicated tool. Write descriptions for tables and columns, and run quality tests in CI. That is all it is, and it will carry you until you hit the wall at around thirty people.
If you are a small or mid-size SaaS
This is the stage for naming data stewards, setting naming conventions and standing up an open-source catalogue such as DataHub. Put a part-time steward on each domain and make table ownership explicit. If you handle personal data, deciding PII tagging and a masking policy at this stage saves a great deal of trouble later.
If you are a large or regulated enterprise
The full kit: Collibra or Alation, a dedicated governance function, and access control down to row and column level. Take the regulatory requirements — GDPR, data-protection law, SOX — as inputs at the start, and put history retention, audit logs and lineage fully in place. This is an area where failing the regulator can mean leaving the business, so it is not the place to economise.
AI decision axes — Governance is a dictionary for the AI
Data catalogs become the search foundation for AI agents
When DataHub or Amundsen are exposed via API, AI agents can autonomously search “which tables relate to revenue” and assemble accurate queries. PDF/Excel glossaries can’t be auto-accessed by AI, making them barriers to AI utilization.
Codified data quality rules verify AI-generated ETL
When data quality rules (NULL check, range validation, referential integrity) are codified in dbt tests or Great Expectations, AI-generated ETL jobs are automatically verified at each run. Without quality gates, defects in AI-generated transformation logic propagate to downstream dashboards and reports undetected.
Pitfalls and forbidden moves
Here are the six most dangerous direct causes of failed audits, regulatory breaches and wrong AI decisions.
| Forbidden move | Why it is bad → what to do instead |
|---|---|
| Believing that installing a tool delivers governance | the catalogue is left alone and rots → pair it with stewards, policy and operations |
| Leaving “tables nobody manages” in place | three years brings thousands of mystery tables → one owner per dataset, without exception |
| Loading personal data into the warehouse without masking | a breach of GDPR and data-protection law → row and column control plus dynamic masking |
| Accumulating data with no retention policy | you cannot answer a GDPR erasure request → design retention periods and automatic deletion |
| Keeping the glossary in PDF or Excel | an AI cannot read it and nobody can search it → use a catalogue reachable by API |
| Not disabling a leaver’s access immediately | it becomes unauthorised access and a leak → automate same-day disabling |
Do not shy away from governance as though it meant “restriction.” Good governance is the foundation that lets data be used safely, and it accelerates use rather than slowing it.
Author’s note - mountains of “ownerless tables” and the EUR 1.2B fine
There’s a story often heard about a mid-size SaaS company where the data-analytics team was actively using dbt and BigQuery, but “didn’t put in governance mechanisms” - and in 3 years, thousands of “tables of unknown origin” piled up. Tables from ex-employees, intermediate tables from experiments, urgent-aggregation tables from half a year ago - none deletable with confidence, leaving only storage cost and confusion.
A more serious case is the May 2023 incident where Meta was fined EUR 1.2B (about JPY 200B) for GDPR violation by transferring EU citizens’ data to the US. The largest record since GDPR took effect, it remains a talking point as a case showing an era where “one storage-location design choice creates fines at the company-survival level.” The early-2017 large-scale MongoDB ransomware case (instances with auth-setting forgotten breached in tens of thousands worldwide) is also told as a representative example of “governance absence connecting directly to incidents.”
I myself, in a previous job, saw groups of tables in the state of “unclear what tables, but can’t confirm if it’s OK to delete,” and felt how governance absence quietly produces debt. Both leave the common lesson that “putting in tools alone doesn’t protect.” The reality that without the trinity of institution, people, and technology, the platform actually becomes liability - is told as a case appearing regardless of scale.
Governance is a trinity of tools, institution, and people. Missing any one and it doesn’t function.
What to decide - what is your project’s answer?
For each of the following, try to articulate your project’s answer in 1-2 sentences. Starting work with these vague always invites later questions like “why did we decide this again?”
- Data catalog (DataHub / Amundsen / Collibra / dbt docs)
- Data stewards (who owns what)
- Quality tests (dbt tests, Great Expectations)
- Access-control method (table/column/row level)
- Personal info handling (masking, retention)
- Data classification (public / internal / confidential / top secret)
- Audit logs (who referenced what when)
Related Articles
Summary
This article covered data governance, including data catalog, metadata, lineage, quality management, stewards, access control, a phased roadmap by scale and regulation, and AI-era governance.
Phased adoption matched to scale, always name stewards, auto-test quality, and curate into AI-readable metadata. That is the practical answer for data governance in 2026.
Next time we’ll start a new category (Security Architecture).
Back to series TOC -> ‘Architecture Crash Course for the Generative-AI Era’: How to Read This Book
I hope you’ll read the next article as well.
Also popular with readers
📚 Series: Architecture Crash Course for the Generative-AI Era (51/95)