User guide

Governance

Measuring what your tables hold, finding the columns that carry personal data, deciding what each one is, and keeping that up to date as tables change.

Governance works in three stages. Profiling measures each column. Classification proposes what personal data each column holds. Review is where a person accepts or rejects each proposal. Nothing changes who can see what until a proposal is accepted. Open it from Catalog → Governance.

Profiling

A profiling run reads every column in its scope — a catalog, a namespace or one table — and records its shape: completeness, distinct counts, lengths, numeric ranges, and which pattern its values follow (an email address, a phone number, a date). It never records the contents. Example values are masked before anything is written, and a column whose pattern is high-risk — a card number, a national ID, an IBAN, a phone number — keeps no example at all, only the name of the pattern.

Start a run from the Governance page, for the catalog chosen in the top bar. A run spends compute, and the page shows its progress: tables profiled, tables requested, and any that failed. The results appear on each table's Profile tab. Quality rules are checked against every new profile.

Classification

Classification proposes which kind of personal data each column holds, using a shared vocabulary of personal-data categories. It works in tiers:

  • Rules run over every profiling run as it lands. They match column names and measured patterns, and propose what they can match with confidence.
  • The model reads the columns the rules could not settle — the ones nobody named after what they hold — and proposes what a person reading the name and the measurements would say. It runs on a model inside the platform; nothing leaves for a third party. It is available only if your deployment has one configured.

A model run goes a table at a time and carries on whether or not the page is open. It skips what the rules already settled confidently, what a reviewer has already decided, and anything it has answered before on the same evidence. Only one model run goes at a time; Stop ends it.

Reviewing proposals

The review queue lists every column with an open proposal, alongside the evidence: its type, its measured pattern and completeness, masked example values, and the column's comment. Flags point out the ones that need judgement — a special category of personal data, a name and pattern that disagree, a model that disagrees with the rules, weak evidence.

  • Accept or Reject each proposed category.
  • Classify as a different category if every proposal is wrong.
  • Nothing here if the column holds no personal data.

Accepted tags are applied to the column. They change what the platform will let out of an export, which is why deciding is for owners and admins.

Watching a scope

A watch keeps a scope up to date without anyone starting a run. Every table in it — including tables created after the watch — is profiled and classified once per schema version: when it first appears, and again when a column is added, dropped or retyped.

A watch runs as a named person, so it reads only what their grants reach. The Watched scopes page shows each watched table's schema version, when it was last profiled and classified, what is waiting its turn, and anything stuck, with the reason. A watch can be paused, resumed or removed.

Describing a table

Classification says what a column holds. Dataset Governance Intake, on a table's page, records why the table exists: its purposes and legal bases, whose data it holds, how long it is kept and how it is protected, and where it lives. Choices arrive pre-ticked from what the table's column tags suggest. Database Infrastructure Governance on a namespace's page does the same for a whole namespace. Both can be exported as JSON-LD.

Who may do what

ToOwnerAdminMemberViewer
See profiles, proposals and decisions✓✓✓✓
Start a profiling run or a model run✓✓✓—
Add and change quality rules✓✓✓—
Accept, reject and reclassify✓✓——
Add, pause and remove watches✓✓——