Data Architecture

[Data Architecture] Data Governance

[Data Architecture] Data Governance

About this article

As the seventh installment of the “Data Architecture” category in the series “Architecture Crash Course for the Generative-AI Era,” this article explains data governance.

A technology-only data platform rots in 3 years. This article covers the components of governance - data catalog, metadata, lineage, quality management, data stewards, and access control - alongside an introduction roadmap by scale and regulation, and the structure where in the AI era governance functions as a dictionary for AI.

Before you read this

This article uses a good deal of vocabulary from around databases. If that is unfamiliar, reading the primer "Database Basics" first makes it far easier to follow. You can also look anything up in the glossary as you read.

What is data governance in the first place

Four Pillars of Data Governance

Data governance is “the mechanism for establishing and continuously enforcing rules on who can use what data and how across the entire company.”

Imagine a condo association. If residents dump trash whenever they want, leave personal belongings in shared areas, and copy keys freely, the whole building falls apart. It’s the management regulations (governance) and the building manager (data steward) that maintain order. Data is the same - without rules for definitions, naming, access permissions, and quality standards, different departments produce different numbers, personal information leaks, and AI learns from bad data. That’s the collapse governance prevents.

Why data governance is needed

Build a data platform and leave it alone, and over time it accumulates “tables nobody knows who made” and “columns with the same name and different definitions” until the platform itself cannot be trusted. Once departments start using data on their own separate definitions, the revenue figure differs by department, and feeding data with unclear definitions to an AI sharply raises the risk of a wrong decision.

The weight of regulation has also changed by an order of magnitude. In May 2023 Meta was fined EUR 1.2 billion under the GDPR for transferring EU citizens’ data to the United States. GDPR fines run up to 4 percent of annual global turnover — this is an era in which “where the data is stored” alone can produce a penalty at the level of company survival.

The main components

To implement data governance, combine the following elements. Any one alone is insufficient - it functions as a trinity of organization, institution, and technology.

Components of Data Governance Like a condo association. Functions as a trinity of organization, policies, and technology Organization Data Steward Person responsible for each dataset Data Owner Final approval authority Data Committee Company-wide policy creation & review Governance doesn't work without people Policies & Rules Policies Retention period, encryption, disposal rules Quality Standards Rules for detecting definition violations, gaps, duplicates Access control rules Permissions for who can see what Without rules, each department has different numbers Technology Data Catalog DataHub / dbt docs / Collibra Metadata + lineage Visualize definitions, owners, and transformation flows Automated quality checks dbt tests / Great Expectations Without automated tool verification, it becomes superficial Functions as a trinity A data platform built on tech alone rots in 3 years. Systems of people and policies must sustain it
ElementRole
Data catalogCatalog of where what is
Metadata managementEach data’s definition, owner, update frequency
LineageVisualization of data transformations and flow
Quality managementDetection of definition violations, missing data, duplicates
Access controlPermissions on who can see what
Data stewardsPerson responsible for each data
PolicyRules for retention period, encryption, disposal

Catalogue and metadata — getting rid of “where is this data?”

A data catalog is the organization’s catalog of all data, with the mechanism of central searchability for “where, what data, with what definition, managed by whom.” It originated when Google built an internal mechanism to “search data like Google Search,” and today commercial and OSS tools are provided by various vendors.

Without a catalog, analysts ask people every time “where is this data?” and “what is this number?”, dropping data utilization speed to 1/10.

ProsCons
Faster data discoveryHeavy initial metadata setup
Shorter new-hire onboardingInformation rots if neglected
Becomes evidence for auditsTool fees, operational cost
Premise for AI integrationDoesn’t function without a steward regime

At small scale, dbt docs is enough. Consider dedicated tools from mid-scale onward.

There are several catalogue tools to pick from.

ToolWhen to choose
dbt docsAlready on dbt, analytics-model-centric. Lightweight and free
DataHub (LinkedIn OSS)Mid-to-large, want to grow OSS while customizing
Amundsen (Lyft OSS)Lighter and simpler than DataHub. Want to lower the entry bar
Apache AtlasLegacy DWH environment centered on Hadoop/Hive
CollibraEnterprise, want to seriously build a governance regime
AlationWant AI features (natural-language search etc.)

DataHub is becoming the OSS de facto. Commercial-side has the two giants Collibra (full-feature) and Alation (strong AI), with investment scales for large enterprises.

Metadata management sits alongside it.

The contents of the data catalog are metadata (data about data). For each table/column, record “what it’s for,” “how to use it,” and “who to ask,” so users don’t get lost.

Metadata typeContent
Business metadataDefinition, meaning, business glossary
Technical metadataSchema, types, constraints, indexes
Operational metadataUpdate frequency, SLA, job history
Owner infoData steward, contact
Quality metadataTest results, anomaly-detection scores

Metadata needs both auto-collection (catalog tool scans) and manual entry (steward writes), and full automation is impossible.

Lineage makes the flow visible.

Data lineage is the mechanism for tracking where data came from, how it’s transformed, and where it flows. The DAG auto-generated by dbt is one form of lineage, visualizing what columns of which business DB this aggregated result is derived from.

Use caseContent
Impact investigation”What breaks if I change this column?”
Cause tracking”Where did this dashboard’s number go wrong?”
Audit response”Where is personal data flowing to?”
Data retirement”Who’s using this table?”

Without lineage in place, you can’t delete old tables, and mystery tables pile up.

Quality management and stewards — the technical and human halves

Data quality is measured along 6 viewpoints. Auto-testing these and validating during pipeline execution is the modern best practice. dbt’s tests feature and Great Expectations are used.

ViewpointMeaningExample
CompletenessNo missing dataNo NULL in required items
UniquenessNo duplicatesUser IDs not duplicated
AccuracyMatches realitySales amounts match actuals
ConsistencyNo logical contradictionsend_date > start_date
TimelinessUpdates aren’t laggingDaily data updated daily
Referential integrityForeign keys existUser ID exists

Data without guaranteed quality becomes the worst risk of misleading executive decisions.

The human half is the data steward.

A data steward is a human role taking responsibility for each data set. The organization needs people deciding the things technology alone can’t solve - “what’s this data’s definition?” and “how may it be used?”

RoleResponsibility scope
Business stewardManage definition, use, business terms
Technical stewardSchema, pipeline, quality
Data ownerFinal approval, scope of disclosure
Data custodianDaily operations, access grants

Stewards are assigned at least one per data set to clarify the locus of responsibility. The basis of governance is not leaving “tables no one manages” alone.

Access control — down to row and column level

Data containing personal/confidential information is strictly managed for who can access. Mere “table-level permissions” aren’t enough - row-level and column-level access control becomes necessary.

Control levelContent
Table levelRead/write permission per table
Column levelSalary column visible only to HR
Row levelSee only your department’s data
Dynamic maskingMask PII (Personally Identifiable Information) at query time
Retention periodAuto-delete after N days

BigQuery and Snowflake natively support column-level and row-level policies, and SQL-side automatic filtering removes the need for app-side individual implementation.

How to choose — a phased roadmap by scale and regulation

Going “large-scale straight away” stalls operations, so growing it in stages is the realistic path.

PhaseOrganisation sizeWhat you introduce
1. Minimumup to 30 peopledbt docs plus dbt tests
2. Basicup to 300 people+ named data stewards, naming conventions
3. Mid-sizeup to 3,000 people+ DataHub or Amundsen, row- and column-level access control
4. Enterprise3,000 and up+ Collibra or Alation, a dedicated governance function
5. Regulated or listedany sizehistory retention, audit logs, full lineage

Regulatory requirements go in at the start. Personal data in the EU falls under GDPR (right to erasure, consent management), healthcare under HIPAA, payments under PCI DSS, listed companies under SOX. In heavily regulated industries governance investment is not optional, and failing to meet it can mean withdrawing from the business. Requirements also rise with how far you take AI: catalogue and consistent naming for BI alone, organised metadata for Text-to-SQL, lineage and quality tests for RAG and agents, and audit logs and explainability for autonomous decisions.

Three scenarios

If you are building solo or at a startup

Starting with dbt docs and dbt tests alone is enough. You need neither a dedicated person nor a dedicated tool. Write descriptions for tables and columns, and run quality tests in CI. That is all it is, and it will carry you until you hit the wall at around thirty people.

Personal / Startup: Ship in One Month Is Correcten.senkohome.com/arch-intro-case-startup/

If you are a small or mid-size SaaS

This is the stage for naming data stewards, setting naming conventions and standing up an open-source catalogue such as DataHub. Put a part-time steward on each domain and make table ownership explicit. If you handle personal data, deciding PII tagging and a masking policy at this stage saves a great deal of trouble later.

Small-Mid SaaS - Lean on Managed and Run with Few Peopleen.senkohome.com/arch-intro-case-saas/

If you are a large or regulated enterprise

The full kit: Collibra or Alation, a dedicated governance function, and access control down to row and column level. Take the regulatory requirements — GDPR, data-protection law, SOX — as inputs at the start, and put history retention, audit logs and lineage fully in place. This is an area where failing the regulator can mean leaving the business, so it is not the place to economise.

Large-Enterprise Core: Design That Holds Up for Yearsen.senkohome.com/arch-intro-case-enterprise/

AI decision axes — Governance is a dictionary for the AI

Data catalogs become the search foundation for AI agents

When DataHub or Amundsen are exposed via API, AI agents can autonomously search “which tables relate to revenue” and assemble accurate queries. PDF/Excel glossaries can’t be auto-accessed by AI, making them barriers to AI utilization.

Codified data quality rules verify AI-generated ETL

When data quality rules (NULL check, range validation, referential integrity) are codified in dbt tests or Great Expectations, AI-generated ETL jobs are automatically verified at each run. Without quality gates, defects in AI-generated transformation logic propagate to downstream dashboards and reports undetected.

Pitfalls and forbidden moves

Here are the six most dangerous direct causes of failed audits, regulatory breaches and wrong AI decisions.

Forbidden moveWhy it is bad → what to do instead
Believing that installing a tool delivers governancethe catalogue is left alone and rots → pair it with stewards, policy and operations
Leaving “tables nobody manages” in placethree years brings thousands of mystery tables → one owner per dataset, without exception
Loading personal data into the warehouse without maskinga breach of GDPR and data-protection law → row and column control plus dynamic masking
Accumulating data with no retention policyyou cannot answer a GDPR erasure request → design retention periods and automatic deletion
Keeping the glossary in PDF or Excelan AI cannot read it and nobody can search it → use a catalogue reachable by API
Not disabling a leaver’s access immediatelyit becomes unauthorised access and a leak → automate same-day disabling

Do not shy away from governance as though it meant “restriction.” Good governance is the foundation that lets data be used safely, and it accelerates use rather than slowing it.

Author’s note - mountains of “ownerless tables” and the EUR 1.2B fine

There’s a story often heard about a mid-size SaaS company where the data-analytics team was actively using dbt and BigQuery, but “didn’t put in governance mechanisms” - and in 3 years, thousands of “tables of unknown origin” piled up. Tables from ex-employees, intermediate tables from experiments, urgent-aggregation tables from half a year ago - none deletable with confidence, leaving only storage cost and confusion.

A more serious case is the May 2023 incident where Meta was fined EUR 1.2B (about JPY 200B) for GDPR violation by transferring EU citizens’ data to the US. The largest record since GDPR took effect, it remains a talking point as a case showing an era where “one storage-location design choice creates fines at the company-survival level.” The early-2017 large-scale MongoDB ransomware case (instances with auth-setting forgotten breached in tens of thousands worldwide) is also told as a representative example of “governance absence connecting directly to incidents.”

I myself, in a previous job, saw groups of tables in the state of “unclear what tables, but can’t confirm if it’s OK to delete,” and felt how governance absence quietly produces debt. Both leave the common lesson that “putting in tools alone doesn’t protect.” The reality that without the trinity of institution, people, and technology, the platform actually becomes liability - is told as a case appearing regardless of scale.

Governance is a trinity of tools, institution, and people. Missing any one and it doesn’t function.

What to decide - what is your project’s answer?

For each of the following, try to articulate your project’s answer in 1-2 sentences. Starting work with these vague always invites later questions like “why did we decide this again?”

  • Data catalog (DataHub / Amundsen / Collibra / dbt docs)
  • Data stewards (who owns what)
  • Quality tests (dbt tests, Great Expectations)
  • Access-control method (table/column/row level)
  • Personal info handling (masking, retention)
  • Data classification (public / internal / confidential / top secret)
  • Audit logs (who referenced what when)

Summary

This article covered data governance, including data catalog, metadata, lineage, quality management, stewards, access control, a phased roadmap by scale and regulation, and AI-era governance.

Phased adoption matched to scale, always name stewards, auto-test quality, and curate into AI-readable metadata. That is the practical answer for data governance in 2026.

Next time we’ll start a new category (Security Architecture).

Back to series TOC -> ‘Architecture Crash Course for the Generative-AI Era’: How to Read This Book

I hope you’ll read the next article as well.