MDM Studio — User Guide
DESIGN · MAP · GOVERN · TRUST
MDM Studio — User Guide
Adaptive Canvas · Enterprise Master Data Management Release v1.2.0 · DESIGN · MAP · GOVERN · TRUST
About this guide
This guide is a complete reference for MDM Studio, the Adaptive Canvas platform for enterprise Master Data Management. It covers the full journey — from signing in and designing a model, through connecting and mapping sources, mastering golden records, governing and measuring quality, to administering tenants, users and sensitive‑data access.
It is written for every role that uses the platform: data stewards and analysts who work with data day to day, developers and data owners who build and configure the model, and administrators who run the platform. Where a task needs a particular role or permission, the guide says so.
Conventions used in this guide
- Bold marks screen labels, buttons and menu names exactly as they appear in the app.
- Italic marks the names of concepts (a golden record, a tenant).
- A Role note at the start of a procedure tells you the minimum role required.
- Steps are numbered; anything you can safely skip is called out.
Product version: This guide describes MDM Studio v1.2.0. The version is shown in the About dialog (open it from the header). MDM Studio ships as a single versioned release covering the web studio and the backend together.
1. Introduction to MDM Studio
1.1 What MDM Studio is
MDM Studio unifies master data — customer, product, party and any other domain — that is scattered across your source systems into governed golden records your whole organisation can trust. It runs end‑to‑end on the database platform you already operate: Microsoft SQL Server, PostgreSQL, MySQL or Oracle.
The platform is organised around four disciplines, which also name its method:
- DESIGN — model your master‑data estate: domains, entities, a governed data dictionary, relationships and ERDs, versioned like software.
- MAP — connect your sources and map their fields, stage by stage, into your model; run scheduled ingestion pipelines.
- GOVERN — control change with review‑and‑approval workflows, policies, role‑based access, sensitive‑data clearances and a full audit trail.
- TRUST — match, merge and survive records into golden records with complete cross‑reference lineage, quality scoring and controlled write‑back.
1.2 Key concepts
Understanding a few core terms makes the rest of the platform intuitive.
| Term | What it means |
|---|---|
| Hub / engine | The database engine that stores the master‑data hub — SQL Server, PostgreSQL, MySQL or Oracle. One canonical model runs natively on each; you choose the engine per session at sign‑in. |
| Tenant | An isolated workspace. Everything you see and change in a session belongs to one tenant; data is never shared across tenants. Tenants have their own branding, users and models. |
| Model | A version/environment of the design — domains, entities, dictionary, mappings, jobs and survivorship rules. A model is a full, independent snapshot; you can clone, export/import and check out models. |
| Domain | A subject area of master data (e.g. Customer, Product). Domains group entities and carry a steward and owner. |
| Entity | A master‑data object within a domain (e.g. Customer, Address). Entities hold attributes and produce golden records. |
| Attribute | A field of an entity, defined in the Data Dictionary with type, keys, classification and sensitivity. |
| Golden record | The single, trusted, mastered version of a real‑world thing, produced by matching and survivorship, with cross‑reference (xref) lineage back to every contributing source. |
| Pipeline stage | One step in the flow from source to golden record: Extract → Land → Stage → Standardize → Match → Survive → Publish. Each stage reads the previous stage's output; a hoverable stage glossary in the app explains each one. |
| Source system / connection | A registered origin of data and the credentialed connection used to read it. Connections are shared across engines in one global registry. |
| Reference data (RDM) | Governed code sets (e.g. country codes) with immutable releases, served to downstream systems through a read‑only API. |
1.3 The platform at a glance
MDM Studio's web studio is divided into ten hubs, shown down the left sidebar. Each hub groups its tools into tabs, and the tabs are themselves grouped into named stages of work.
| Hub | Tab groups and tools |
|---|---|
| Home | Overview (the Dashboard) · My Work — your unified task inbox (§2.9) |
| Modelling Studio | Plan — As‑Is Assessment, Business Glossary · Design — Data Modelling (ERD), Model Builder, Domains & Entities, Hierarchies · Document — Data Dictionary, Catalog, External Assets |
| Integration Hub | Configure — Sources & Connections, Mapping Designer, Standardization · Operate — Jobs & Runs, Orchestration |
| Quality | Assess — Data Profiling, Placeholder Advisor, Source DQ, Source Comparison · Measure — DQ Rules, Data Quality · Remediate — Exceptions, Campaigns |
| Mastering | Configure — Mastering Rules · Steward — Golden Records, Author Records · Integrate — Real‑time API, Write‑Back · Trace — Pipeline / Lineage, Record Timeline |
| Reference Data | Overview · RDM Explorer, RDM Duplicates, RDM Standards, RDM Change Requests, RDM Dashboard, RDM Distribution |
| Governance | Policies, Change Requests, Approval Workflows |
| Analytics | Consume — GR Explorer, Report Center, Dashboards · Build — Report Studio · Operations — Pipeline Health |
| Administration | Identity — Users, Roles & Permissions, Access Control, Active Directory, SCIM Provisioning · Security — Security Posture, Secrets Vault, Data Access · Monitoring — Audit Trail, Active Sessions, Performance, Scale Readiness · Storage — Backup & Restore, Attachment Storage · Reference data — RDM Serving, RDM Integration, Catalog Integration · Platform — Tenants, Deployment Models · Account — My Account |
| Help | The product guides, and an assistant that answers questions from them (§2.7) |
Why the Reference Data tools all begin with "RDM". Explorer, Dashboard and Change Requests each already exist elsewhere in the studio as master‑data surfaces. The Reference Data hub prefixes its own with RDM so that a hub tile, a browser tab title and a result in the command palette (§2.2) name one thing and one thing only. The prefix is a label: the addresses behind the tools are unchanged, so existing links, bookmarks and deep links still resolve.
Your sidebar may be shorter than this. The sidebar is shaped by your role. If you are a Steward, Data Owner or Analyst, the Modelling Studio, Integration Hub and Administration hubs are tucked under a single collapsed Engineering disclosure, so the hubs you use daily sit at the top. Expand Engineering to reach them. Administrators, developers and viewers see the ten hubs flat. Separately, any tab your role may not open is hidden from both the tab strip and the hub landing page — so a shorter list is normal, not a fault.
Two tools are not in the sidebar at all. Messages (§2.11) opens from the header, and Licensing (§10.19) opens from the profile menu. A handful of screens — the Dashboard Builder, the Report Viewer, the relationships editor — are reached by opening the thing they edit rather than from a tab.
The process map. Home and every hub landing show a persistent pipeline strip — Sources → Ingest → Profile → Standardize → Match → Survive → Publish/Golden, plus the operate loop DQ rules → Exceptions → Campaigns → Write‑back. Every node is a deep link to the right tool, so the strip doubles as a map of the method: build left to right, then operate the loop.
When a hub has more tools than the bar can hold. The tab strip measures itself against your window. Where everything fits, you get the tools in their natural order. Where a hub is large enough that tools would have to be hidden — Administration is the one that reaches this on most screens — the bar switches to carrying the groups instead, each with the number of tools in it, and the selected group's tools appear on a slim row beneath. Clicking a group reveals its tools; it does not navigate, so you can look without leaving the page you are working in. Navigating from the row below snaps the bar back to the group you are actually in. On a wide enough window a hub returns to the flat bar by itself — the shape follows the space, not a setting.
Every tool has its own address, so any screen can be bookmarked or shared. Older links still work — the studio redirects roughly three dozen retired addresses to their current home.
The rest of this guide follows these hubs in order, after getting you signed in and oriented.
2. Getting started
2.1 Signing in
Open MDM Studio in your browser (for the reference deployment, https://mdm.adaptivecanvas.co.za). The sign‑in page asks for your username (or email) and password.
MDM Studio supports two sign‑in methods, chosen per user by an administrator:
- Local — your password is stored securely (hashed) in MDM Studio.
- Active Directory (LDAP) — your password is verified against your organisation's directory. You sign in exactly the same way; the platform routes the check to the directory behind the scenes.
Choosing a database engine. If your deployment offers more than one engine, you can pick which hub engine to work against at sign‑in. Your session stays pinned to that engine.
Multi‑tenant sign‑in. If your account belongs to more than one tenant, after your password is verified you are asked to pick which tenant to enter — each tenant button carries that tenant's own badge colours and logo. The list only ever shows tenants you have access to — tenant names are never visible before you authenticate. Single‑tenant users go straight in. A short confirmation, styled in the tenant's branding, shows which tenant you signed into.
Where you land. MDM Studio remembers where you were: signing back in returns you to the page you last worked on (first‑time sign‑ins land on Home).
Idle sign‑out. For security, MDM Studio signs you out automatically after a period of inactivity (10 minutes by default; configurable per deployment). Ten seconds before it happens a warning dialog appears with a live countdown — click Stay signed in to continue, or Sign out now to leave immediately. If you do nothing, the session ends and you return to the sign‑in page.
2.2 The workspace
Once signed in you land on Home (the Dashboard). The workspace has three regions.
The header (top bar), left to right:
- Tenant chip — the tenant you are working in, with its logo/branding. Everything in the session is isolated to this tenant.
- Model switcher — the active working model. Click it to switch models or to open Deployment Models. Model‑specific tools require a working model to be selected.
- Checkout control — check the active model out for exclusive editing, or request control if someone else holds it (see §2.4).
- Review‑bypass pill — appears only while governance review is suspended for the active model. When the bypass is on, an amber Review bypass ON pill sits here (with a Turn off action) and the whole header takes on an amber theme, so nobody works in a governance‑suspended model without noticing. When the bypass is off — normal governance protocols — there is no pill and the header is plain: the chrome speaks only when something has been switched off. Turning the bypass on is done from Administration → Deployment Models, not from the header. See §7.3.
- Search (⌘K / Ctrl‑K) — the command palette, to jump anywhere quickly.
- Messages — direct messages and threads with colleagues (§2.11).
- Notifications, About, and your profile menu — the profile menu holds My Account, tenant info, Licensing (administrators) and Sign out.
The browser tab shows Adaptive Canvas – \<tenant code>, and tenants with a logo get it as the tab's icon — so multiple tenant sessions are easy to tell apart in the tab strip.
The sidebar lists the hubs your role can open. Click a hub; its tabs appear across the top of the workspace, grouped by stage of work. The sidebar can be collapsed to icons. Stewards, data owners and analysts find Modelling Studio, Integration Hub and Administration inside the collapsed Engineering disclosure at the foot of the list (§1.3).
The workspace area shows the selected tab. Model‑specific pages display a prompt to select or create a working model if none is active.
Moving around. Every hub tab strip includes a ← Back button, and the Backspace key steps back through the pages you visited (both hide on your first page, so you can never back out into the sign‑in screen). Navigation is instant: pages you've visited are served from a smart cache and refreshed automatically after any edit, model switch or sign‑in — use the emerald Refresh button on any page to force a reload at any time.
2.3 Models and the working model
Most design and mastering work happens inside a model. A model is a complete, independent snapshot of the design.
- Use the model switcher in the header to choose your active working model.
- Manage models in Administration → Deployment Models: create a new model, clone an existing one, import/export a model as XML, or archive one.
- When you create a model (or change its target database), MDM Studio provisions the target database and all pipeline schemas automatically on SQL Server, PostgreSQL and MySQL — no manual database creation needed. Any provisioning problem is reported but never blocks creating the model. (On Oracle, provisioning is skipped with an explanation.)
- The special base store holds only shared, cross‑model items (connections and managed/reference values) and cannot be edited directly as a working model.
2.4 Model checkout (exclusive editing)
To prevent two people editing the same model at once, MDM Studio uses checkout.
Role: Any role that can edit the model.
- Check out the active model from the header to hold the exclusive edit lock.
- Check in when you're done, reviewing what you changed on the way out.
- If someone else holds the lock, you'll see who has it, how long they have held it, and can Request control; they get a prompt to hand it over.
2.4.1 What the lock does and does not cover
The checkout is a lock over the model's definition — domains, entities, attributes, relationships, mappings, match and survivorship rules, standardization, jobs and glossary terms. Two people cannot change those at once, and that is the whole purpose of the lock.
It is not a lock over the hub. While someone holds a checkout, everyone else can still run ingestion jobs, review match pairs, resolve data‑quality exceptions, edit and author golden records, publish reference‑data releases and approve governance requests. None of that changes the model's definition, and none of it is arbitrated by the lock.
2.4.2 Reviewing changes at check‑in
Checking out takes a snapshot of the model's definition. At check‑in the studio compares the model against that snapshot and lists what changed, grouped by area, with a plain summary per change and the field names behind it. Every change is ticked by default: leave it ticked to commit, untick it to revert it to its exact pre‑checkout value. Reverts run in one transaction — if any part fails, none of it is applied and the lock is kept.
Three things the review will tell you, because each of them is a case where a silent answer would be wrong:
- Changes approved through governance are marked. Where a change request was approved and activated while you held the checkout, that change carries a badge naming the request and its approver. Reverting one is still permitted — but it asks you to confirm that you mean to overrule an approval, and the fact is recorded in the audit trail against your name.
- Areas it could not compare are named. If the snapshot could not read part of the model, the review says which areas, and those changes are checked in as they stand rather than being listed as new.
- A failed review is not "nothing changed". If the change list cannot be loaded, the dialog says so and warns that checking in will commit everything unreviewed — it does not tell you there was nothing to review.
An administrator can release a checkout held by someone who is unavailable. Releasing drops the lock without taking the model over, and keeps the absent holder's snapshot, so their pending changes are still reviewable at the next check‑in. It is recorded in the audit trail.
2.5 Roles at a glance
MDM Studio has six roles, from least to most privileged. Exact permissions are configurable per deployment (see §10.3), but the defaults are:
| Role | Typical use |
|---|---|
| Viewer | Read‑only access to permitted areas. |
| Analyst | Explore data, run reports, work exceptions assigned to them. |
| Steward | Day‑to‑day data stewardship — modelling, mapping, matching, quality remediation. |
| Data Owner | Accountable owner for domains; user management within their remit. |
| Developer | Build and configure — models, sources, mappings, custom report sources. |
| Administrator (ADMIN) | Full platform administration — users, roles, tenants, sensitive‑data access, Active Directory, deployment. |
A separate concept, platform superuser, governs cross‑tenant administration (managing tenants). Superusers are assigned by configuration and/or granted in the app; see §10.1.
2.6 My Account
Role: Every role can manage their own account.
Open My Account from the profile menu (Administration → My Account). Here you can:
- Change your display name and email.
- Choose an avatar from the built‑in gallery, or upload an image (up to the configured size limit, 5 MB by default).
- Change your password (local accounts only — Active Directory accounts manage their password in the directory).
- If you are a tenant administrator, customise your tenant's badge branding (logo, colours, font). Reset style clears and saves the styling immediately.
- See the tenants your account can access.
- Choose your theme — light or dark. The choice is stored on your account rather than in the browser, so it follows you to any machine, and it is applied before the first pixel is painted rather than flashing the wrong theme first.
- Set your own value for the shared display settings below.
2.6.1 Display settings, and which tier answered
Some settings shape how the product reads rather than what it does. These resolve through three tiers, in order: your own preference, then the model default, then the system default. Whichever tier answered is shown beside the control, so "why is mine different from yours" is a question the screen answers.
- Filter dropdown threshold — the number of distinct values below which a report or dashboard filter renders as a picker rather than a text box (§8.2.3). Setting it to 0 turns pickers off entirely.
Your own value is yours to change. Changing the model default, which is what everyone without a personal value gets, is a governed action restricted to governance roles and written to the audit trail.
2.7 Help, guides and the AI assistant
Role: every role.
The Help hub (sidebar, or the ? shortcut in the header) puts the product documentation one click away:
- Guides — this User Guide and the step‑by‑step Domain Walkthrough open right inside the studio, with a navigable sidebar.
- Ask the assistant — type a question in plain language and the assistant answers from the documentation, quoting and deep‑linking to the exact section it drew from. If no AI model is configured for the deployment, the assistant falls back to ranked search over the same content — help is always available. §12.6 sets out exactly which parts of the product use a language model and what is and is not sent to it.
2.8 Your first model (first‑run experience)
On a fresh tenant with no working model, Home becomes a create your first model panel that takes you straight to the point. Once a model exists:
- A Getting started card on Home tracks ten setup steps — from creating a domain to publishing golden records — showing live done/to‑do state, each step deep‑linked to the right tool.
- Hub landing pages list their tools as a numbered "Do these in order" flow, and empty inventory tiles turn into setup prompts that take you to the tool that fills them.
- The pipeline strip (§1.3) is always one click from wherever you need to go next.
2.9 My Work
Role: every role; the queue is filtered to what is yours.
Home → My Work is one inbox for everything waiting on you, so you do not have to tour half the product to find out what today looks like. It reads through ten kinds of work:
| Work type | What it is |
|---|---|
| Change request | A governed change awaiting your review (§7.2). |
| Exception | A data‑quality failure assigned to you (§6.3). |
| Match review | A candidate pair the matcher scored into the review band (§5.1). |
| Campaign | A remediation campaign you own (§6.5). |
| Real‑time review | An inbound record the real‑time API decided needs a human (§5.7). |
| Workflow approval | A step in an approval workflow assigned to your role (§7.6). |
| Survivorship conflict | Sources that disagree on a golden value and cannot be resolved by rule (§5.1). |
| Comparison conflict | A difference found by Source Comparison (§6.9). |
| Suspension expiry | A rule suspension about to lapse — a deadline the system will honour on its own (§6.6). |
| Noisy‑rule proposal | A suggestion that a rule be suspended — a question with no deadline (§6.6). |
Items are ranked by SLA, so the thing that will breach first is at the top. Every item carries a deep link to the exact record, rule or request behind it — clicking takes you to the destination page with that item already selected. If a task is genuinely not yours today, snooze it: the snooze is personal and hides the item from your queue only.
The queue is grouped, and the groups follow the collapse rule described in §2.10 — the first time you open it everything is folded, and after that it restores exactly how you left it.
The Home overview follows your role too. Stewards and data owners land on their work queue, analysts land on analytics, everyone else lands on the Dashboard.
2.10 Working in grids and trees
Two behaviours are shared by most of the studio and worth learning once.
Right‑click a row. Twenty‑four grids across the product carry a row menu — golden records, exceptions, the data dictionary, glossary, mapping sets, match and survivorship rules, DQ rules, campaigns, sources and connections, users, sessions, models, the audit trail, the catalog, GR Explorer and more. It is the studio's own menu, not the browser's, so it looks and behaves the same everywhere. Escape, a click outside, a scroll or a window resize closes it, and near the edge of the screen it flips rather than running off.
The menu knows which column you right‑clicked, which is what makes Filter by this value land on the right field.
- On golden‑record grids you get: view this record's exceptions · open the record · trace it through the pipeline · open its lineage web · show sample values · filter by this value · copy the golden key.
- On metadata grids you get: open · open where it lives (jump to the page that actually owns the object, not a read‑only echo of it) · the page's own extras · filter · copy · delete.
Two rules make the menu safe to use:
- A menu item is never a shorter path than the page. Anything destructive opens the same confirmation dialog you would get from the page's own button. Deleting from a menu is not a quicker way to delete.
- A greyed item tells you why. Disabled entries carry a written reason ("this rule is in use by 3 rule sets"), never a silent grey.
Some menus are deliberately thin. The Deployment Models menu is navigation only — nothing on it calls the platform. The Audit Trail menu has no delete, because audit records cannot be deleted; that is a property of the audit trail, not an omission.
Trees and groups start collapsed. The first time you open a grouped page — My Work, Jobs & Runs, Job Groups, Change Requests, the rules browser, suggested campaigns, the Report Studio report list — everything is folded, so you see the shape before the detail. From then on the page restores whatever you last left open, including "everything", which is why an expand all sticks between visits. Searching the rules browser reveals matches inside folded groups without permanently unfolding them.
Two details that follow from the rule rather than fighting it: a folded group always shows its count, so you know how much is behind it before you open it; and a group is never allowed to hide the thing you are working in — opening a report in Report Studio unfolds its group, once, and you can fold it again afterwards and have that stick.
2.11 Notifications and Messages
The bell in the header collects notifications — a request routed to you, a run that finished, a clearance granted. Notifications arrive over a live push channel, so the bell updates without a refresh.
Messages (also in the header, not the sidebar) is person‑to‑person messaging inside the platform: threads, unread counts and @mentions. A message can carry the context of a work item, so "can you look at this?" arrives with the exception or change request attached rather than pasted as a description.
3. Modelling Studio
The Modelling Studio is where you design the shape of your master data. Its tabs are grouped into three stages: Plan (As‑Is Assessment, Business Glossary), Design (Data Modelling, Model Builder, Domains & Entities, Hierarchies) and Document (Data Dictionary, Catalog). If you are starting a master‑data programme rather than extending one, begin with the As‑Is Assessment (§3.6).
Role: modelling generally requires Steward, Developer, Data Owner or Admin.
3.1 Data Modelling (ERD)
The Data Modelling tab is a visual ERD canvas. Drag entities onto the canvas, add attributes, and draw relationships between them. You can also pull a standard industry domain into your model — this creates the domain, its entities, attributes (with types, keys and PII flags) and relationships, ready to refine.
Working with relationships.
- Every relationship line carries an R# key badge, matched by a relationships legend — a floating panel you can move, resize and toggle on the canvas; its position is remembered per user. Clicking or hovering a line highlights its legend row, and vice versa.
- Right‑click a relationship to edit it, show or hide its label, pick a line colour, change its type (1:1, 1:M, M:1, M:N) or delete it. Labels are movable, resizable chips with wrapping text; all styling is saved with the model.
Containment on the canvas. Value‑hierarchy containment (§3.10) is drawn on the ERD alongside the foreign‑key relationships, and it is deliberately styled so the two can never be confused: a declared containment is a violet dashed edge between the two participating columns, a proposed one is a fainter amber dot‑dash. Columns that participate carry small ⊃ / ⊂ chips (amber while only proposed), and the Contains button on the toolbar shows a live count of declared and proposed edges, with an eye toggle to hide them on a crowded canvas (your choice is remembered).
Every containment edge answers a click. A proposed edge is labelled proposed — click to review: clicking opens a review naming the entity, the chain, and which ends rest on inference, and accepting records the containment — only ends still inferred are written; a declared end is never re‑stamped. A declared edge opens a read‑only statement of the recorded fact — "PROVINCE_ZA contains CITY_ZA" — with one button to the Value Hierarchies governance page. Read‑only on purpose: un‑declaring a containment changes what every consumer of the chain means, so that stays behind the governance page's own controls. A level the product cannot name gets no line at all — no edge is drawn to a column it would have to guess.
A code set can be created from the column itself. When you declare containment on a column that names no governed code set, the dialog offers "…or create a new code set from this column's data" — the virtual ref list. It is preview‑then‑commit: the first click writes nothing and states exactly what would happen ("a new DRAFT reference domain will be created and seeded once with N distinct values"), with a sample, the blank count and a declared value cap. The commit re‑runs the preview — it cannot state one thing and do another — and creates the domain as DRAFT: values are trimmed and de‑duplicated case‑insensitively, blanks are counted but never seeded, and a re‑run is additive‑idempotent (existing codes are skipped, never duplicated). A masked column refuses before a single value is read — a code list built from a masked column would be the extraction the mask exists to prevent — and an all‑blank column refuses rather than creating an empty domain. Publishing the new domain stays with the normal reference‑data release flow.
Layout & export.
- Auto‑layout offers two arrangements: Compact (a tight grid of collapsed cards) and Expanded (all columns visible, no overlaps, the canvas grows to fit).
- The Export button produces three formats, chosen in one dialog: a visual PNG of the diagram (cards fully expanded, legend included), a translated SQL DDL script (CREATE TABLE with types, nullability and PK/FK constraints), and a translated Excel workbook (Entities / Attributes / Relationships sheets).
3.2 Model Builder & Schema Importer
The Model Builder tab combines two tools:
- Model Builder — assemble a model quickly from building blocks.
- Schema Importer — introspect a source schema and import its tables and columns as entities and attributes.
The Schema Importer is stage‑aware. Pick a pipeline stage and the importer browses that stage's schemas, tables and previews; post‑Extract stages show the mapping feed from the previous stage rather than pretending everything is an Extract. Add all links columns to existing attributes case‑insensitively (no duplicate attributes minted by casing) and registers the mappings against the selected stage; columns that match no attribute at that stage are reported in an unmatched at this stage banner rather than silently skipped. Every registered mapping can be removed individually, with a confirmation dialog — here and on every other surface that shows mappings. The import queue starts collapsed so the page leads with the schema, not the backlog.
The builder is AI‑assisted end to end: for every staged column it can suggest names, business terms, descriptions, classifications, PII flags, a data‑access (sensitivity) level and a golden‑record eligibility profile with a weighted score — so sensitive fields and golden‑record candidates are flagged from the very start. The staging grid shows an Access selector and a GR score chip per column.
Where no AI model is configured for your deployment these suggestions fall back to deterministic heuristics and the builder keeps working; each suggestion reports which engine produced it (§12.6).
While staging you also get live previews straight from the source — per‑column distinct values and a top‑N rows preview per table — so you can sanity‑check the data before you commit the model.
Other builder capabilities:
- Steward and Owner are required when the builder creates a new domain (reusing an existing domain is exempt). Use Bulk assign to set Domain, Steward and Owner on all — or a ticked subset of — staged tables in one action.
- Hierarchy tier wiring at build time — bind up to four hierarchy tiers to columns per staged table; an AI suggestion profiles your live source data and proposes tier columns (e.g. country → region → city → segment).
3.3 Domains & Entities
The Domains & Entities tab manages the two structural layers:
- Domains — create a domain with a code, name, description, data steward (required), data owner (required) and tags; steward and owner are chosen from simple dropdowns. You can also set a default sensitivity level for everything under the domain (see Chapter 11). Duplicating a domain asks for confirmation and auto‑renames the copied entities "(copy)" so their pipeline tables can't collide; entity names are unique across domains.
- Entities — create entities within a domain, set the golden‑entity flag (whether it produces golden records), optionally override the physical hub table name, and set a default sensitivity level that can raise the domain default. The AI can suggest an entity name and description; duplicating an entity asks for confirmation first.
The attribute editor (identical here and in the Data Dictionary) covers everything an attribute carries: access level, reference domain, default value, validation rule, tags, the eight golden‑record criteria with a live weighted score, AI Suggestions (including AI GR scoring), standardization rule assignment, hierarchy tier binding and live distinct‑value previews from mapped sources. The attribute list shows Access, GR and Ref‑domain columns, and attributes can be drag‑reordered in a Reorder modal — list position becomes the ordinal.
Excel round trip. Export entities and attributes to Excel and re‑import them losslessly: the export includes every field with headers matching the import template, and the template ships a Domains reference sheet. Imports accept sensitivity, default values, validation rules, tags and ordinal positions. Exports ask for confirmation first.
Golden preview. The entity's Physical tab can preview the top published golden rows for the entity — masked according to your clearance, with locked columns indicated.
3.4 Data Dictionary
The Data Dictionary is the authoritative catalogue of every attribute. Role: Steward, Developer, Data Owner or Admin to edit.
For each attribute you define:
- Name, code, data type, size/precision and keys (primary, foreign, alternate).
- Nullability, default value and validation rule.
- Classification — IDENTIFIER, CONTACT, FINANCIAL, DEMOGRAPHIC, OPERATIONAL, REFERENCE or AUDIT.
- PII and Sensitive flags.
- Sensitivity level — Public, Internal, Confidential or Restricted (Chapter 11). Values above a user's clearance are masked at serving time.
- Golden‑record scoring — the weighting that decides whether the attribute participates in the golden record.
- A link to the business glossary term (the semantic layer).
The dictionary's attribute editor has the same full AI assistance as the Entity Manager: suggested descriptions and terms, AI golden‑record scoring, a hierarchy‑tier "does this fit as a tier?" verdict, and standardization‑rule suggestions with one‑click apply. The AI glossary link matches an attribute's business term to the glossary and either links the existing term or drafts a new entry for it. Nothing here is applied without your acceptance, and every suggestion reports the engine behind it (§12.6).
Working in the grid:
- The stat tiles (Total / Primary Keys / PII / Sensitive / Mapped) are click‑to‑filter — click a tile to filter the grid, click again to clear.
- Attributes mapped to several sources show a per‑source preview (distinct values and top rows per source connection).
- Open in Entity Catalog deep‑links from any attribute straight to its entity with the attribute highlighted.
- The export / template / import cover the full attribute shape — sensitivity, reference domain, default value, validation rule and tags round‑trip cleanly.
3.4.1 Display formats
An attribute can carry a display format — a named preset such as a thousands‑separated integer, a currency, a percentage, a date style or a masked identifier. It is set once here, on the attribute, and every surface honours it: grids, report tables, chart data labels, axis ticks, KPI cards and record screens. When you import a schema the importer proposes a format from the source column's type and scale; you accept or change it.
A display format is display only. It changes what a value looks like and never what it is. Filtering, copying, CSV and Excel export, generated SQL, data‑quality rules, mapping specifications and every API payload carry the raw value. If a format made a filter match differently from the underlying data, the two would disagree about the same record — so it does not.
3.5 Business Glossary
The Business Glossary is the plain‑language semantic layer: business terms and definitions that can be linked to dictionary attributes, so technical fields carry their business meaning. Terms move DRAFT → PUBLISHED through an approval step. The AI can propose links from the Data Dictionary (§3.4). Terms linked to a model's attributes and domains are cleaned up with the model when it is reset or deleted.
A glossary term is more than a definition. Each term carries:
- Typed relations to other terms — synonym, related, broader, narrower, see also — so the glossary is a network rather than a flat list.
- Bindings to assets, many‑to‑many: one term can govern many columns and attributes, and one asset can be governed by several terms.
- A sensitivity class — NONE, PERSONAL, FINANCIAL or SPECIAL — used by the catalog and by classification proposals. Where bound terms disagree, the most restrictive class wins.
- A masking policy — the shape a redaction takes when a value is masked (§11.6). This is a closed list, not free text.
- Accountable and Responsible people, which the catalog reads as the A and R of RACI (§7.7).
Where an older glossary held its relationships as a free‑text list, migration is conservative: fragments it cannot resolve to a real term are not guessed at — they are put on a review list for a human.
3.6 As‑Is Assessment
Role: Steward, Data Owner, Developer or Admin.
Modelling Studio → As‑Is Assessment is the structured "where are we now" exercise you run before the build, and revisit each round afterwards. You score a domain across ten dimensions on an L1–L5 maturity scale, with weights, and the tool computes the domain's position.
Around the score sit the things that decide sequencing:
- Organisational readiness and business criticality per domain.
- Dependencies between domains, so the roadmap does not schedule a domain before what it needs.
- A generated roadmap and a portfolio view across all assessed domains.
- Stakeholder governance — each stakeholder carries a RACI letter, a power/interest position and a rung on an engagement ladder, and each planned activity carries its own RACI matrix.
- A sign‑off that locks the round, so a baseline stays a baseline.
The assessment is advisory. Nothing in it gates the pipeline: a domain assessed at L1 can be modelled, mastered and published exactly like any other. Its value is in argument and sequencing, not enforcement.
3.7 Hierarchies
The word hierarchy means three different things in MDM Studio. All three are real, and they are not the same feature: managed hierarchies and hierarchy tiers are described here, and value hierarchies — containment between governed code sets, Gauteng contains Johannesburg — have their own section (§3.10).
1 · Managed hierarchies — trees over your master data. Modelling Studio → Hierarchies builds and maintains named hierarchies over golden records: a product taxonomy, a legal‑entity tree, a cost‑centre rollup. It supports ragged hierarchies (branches of unequal depth) and alternate hierarchies (several views of the same members), node create/edit/move, level definitions, top nodes and roll‑ups. Two checks run on demand: an integrity check (orphans, cycles, members in more than one place where that is not allowed) and a placement check for members that have not been placed at all. A hierarchy can be frozen as an immutable version snapshot, and two snapshots can be diffed and exported to CSV — which is how you answer "what changed in the org tree between March and June".
2 · Hierarchy tiers — a four‑level dimension stamped through the pipeline. Each model names four tiers (for example Company → Division → Department → Team). Per entity you bind each tier to an attribute. From then on the tier values are materialised as four columns that are stamped at every pipeline stage and carried onto the golden record — which is what makes per‑tier data‑quality scorecards and per‑tier reporting possible.
Tiers are configured on the model, and wired in three places: the model‑create form (with tier presets and an advisor that proposes tier levels from a short description of your organisation), the Model Builder at build time (§3.2), and per attribute in Domains & Entities and the Data Dictionary. Accepting an advisor's proposal saves it as a tenant preset for future models.
The Hierarchies & Tiers page. The Hierarchies tab opens the Hierarchies & Tiers workbench, which manages both meanings in one place: per‑entity tier bindings with coverage bars (a null bar means not yet stamped, which is not the same as zero), a Backfill action that stamps already‑published golden records, a tier‑level advisor, and — for hierarchies generated from your data — What changed?, Re‑sync and Refresh measures actions. Re‑sync always previews its changes before applying them, and refresh‑measures reports set, unplaced and non‑numeric counts separately rather than as one blended number.
Node labels — names, not codes. A tree is only navigable if you can read it. Every surface that renders a hierarchy — the Hierarchies workbench, the GR Explorer rail, Data Quality's By‑hierarchy view — resolves each node's label attribute: the attribute whose value is shown as the node's name (the category name, not its code, with the code muted after it). You choose it when creating a hierarchy from a parent relationship (the Node labels from picker), and you can change it later from any surface that shows the tree — the change is saved on the hierarchy and every surface follows.
Hierarchies can be suggested, never assumed. When an entity's published data behaves like a tree — a column that resolves into the entity's own key — but no self‑relationship is declared, the Explorer shows a measured one‑line suggestion with the evidence (how many records resolve, how many do not, and how many roots), and Declare this hierarchy / Dismiss buttons. Declaring goes through normal governance; the platform never infers a hierarchy from column names alone, and a dismissal lasts for the session.
Working the tree. Right‑click a hierarchy for Open, Copy and Delete — delete states the node count and asks for confirmation first. Trees are served as one persisted treeview: expansion survives navigation and sign‑ins, an oversized tree says it was capped rather than pretending the visible slice is everything, and a capped branch is never rendered as a leaf.
Tiers align with standardization. When a tier is bound to an attribute that runs standardization rules, the tier value is stamped from the standardized form — so consumer and CONSUMER group under one node instead of two. If the attribute also masters its cleansed value (§5.1), the tier and the stored golden value agree exactly. If it does not — the rules run but the cleansed value is not published — the tier still groups by the cleansed form, and the surfaces say so: the tiers page marks the binding "standardized for grouping — stored value is raw", and the GR Explorer's tier ladder carries the same note. That divergence is deliberate and temporary; the mastering offer on the tiers page (§5.1) is the way out, making tiers, grids and exports agree. Both stamping paths — the pipeline's Publish and the Backfill button here — apply the same rules through the same engine, so a freshly published record and a backfilled one cannot disagree.
Tier names per domain. A model names its four tiers once, and one vocabulary rarely fits every domain — Customer reads Region → Country → City while Product reads Category → Subcategory → Line. A domain can therefore rename any of the four tiers for its own entities on its edit form (Modelling Studio → Domains). A blank box inherits the model's name for that level, so renaming one tier leaves the other three following the model, and every tier surface labels each level with the name that applies and says whether it came from the domain or the model. Report field labels deliberately keep the model's names — a report can cross domains, so no single domain's vocabulary can stand for all of them.
All four bindings on the entity. The entity edit form (Modelling Studio → Domains & Entities) shows the four tier levels with a dropdown each, so an entity's wiring can be seen and changed in one place rather than one attribute at a time. Each change saves on its own, per level, and never disturbs the other three. A new binding is stamped from the next publish; to apply it to records already published, use Re‑stamp on Hierarchies & Tiers. Editing tier names and bindings needs the same role as the model form (Admin, Steward, Data Owner or Developer).
A caution worth stating. Features that read tiers — Data Quality by tier, the survivorship decision matrix, the Tiers lens in By‑hierarchy — show nothing (or a disabled control with the reason written on it) until tier bindings exist on the entity. An empty per‑tier scorecard almost always means unbound tiers, not missing data.
3.8 Catalog
Role: every role can search and read the catalog; editing a facet is a governed change (below).
Modelling Studio → Catalog is a single searchable index across everything the platform already knows about: connections, tables and columns, domains, entities and attributes, mapping sets, relationships, glossary terms. It is a read model projected over the modules that own that metadata, not a second repository — nothing is authored here that is not authored somewhere else first, which is why it can never disagree with the truth for long.
How assets are addressed. Every asset has a stable name of the form ac:{connection}/{database}/{schema}/{table}[/{column}] for physical assets, ac:entity/{code}[/{attribute}] for model assets, and ac:term/{name} for glossary terms. That string is the join key, and it is what appears in exports and links.
What you can do
- Search across every indexed asset. Results are ranked deterministically: an exact match outranks a name match, which outranks a synonym, a URN, a related term and finally a description hit.
- Read an asset's page — where it lives, its type, its glossary term, its classification, its owners, and (for profiled columns) its statistics.
- Set facets. A facet is a piece of curated metadata on an asset — owner, steward, consulted, informed, domain, classification, tags. Facets are set‑valued: saving replaces the whole set for that facet.
- Coverage — which parts of the estate have an owner, a term, a description; the gaps are the work list.
- Impact analysis — "what breaks if I change this?" — following the asset through mappings, rules and reports.
- Ask — a question box that answers from catalog metadata (see §12.6 for what it does and does not send anywhere).
- Proposals — a steward inbox of suggestions: bind this asset to this term, classify this term, create this term, describe this asset. Nothing is applied without an explicit accept, and a rejection is remembered so the same suggestion does not come back next week.
Classification proposals are made on the governing term, not on each column — accept once and every asset bound to that term inherits it.
Three things to know about how it stays current
- The index is rebuilt, not live. It refreshes when somebody presses Reindex, or on a drain interval if your deployment has configured one. A drain interval puts a ceiling on staleness; it does not make the catalog a live view.
- Five kinds of change do not raise a flag at all. Connections, mapping sets, relationships, attribute sources and source trust scores can change without the catalog knowing it should look again — the status panel names them as sources that do not emit. Reindex after that kind of work.
- Reindex is safe. It rebuilds the module‑derived facets wholesale and never touches a facet a human set. Assets that have vanished from their source module are tombstoned, not deleted, so their history and their facets survive. Where a human facet and a term disagree, the human wins; an approved term beats a draft one.
Editing a facet is the one metadata write that goes through Change Governance (§7.2). That is deliberate: a facet is a statement about ownership and meaning, and it should be reviewable.
Honest limits.
- This is not the "Catalog Integration" page. Catalog (here, in Modelling Studio) is the internal read model. Catalog Integration (Administration, §10.18) publishes your metadata to, and pulls it from, an external catalog such as Atlas, Purview or Collibra. Different products, similar names.
- Catalog search is lexical, not semantic, in a stock deployment. The default index matches shared words and shared word‑shapes ("customer" ≈ "customers") and nothing beyond that — it has no model of meaning, and the platform reports it as non‑semantic in its own status panel. A learned model can be configured by your administrator; until it is, treat search as very good keyword matching.
- Renaming a table or column mints a new address and orphans the facets a person typed onto the old one. Re‑attach them after a rename.
- Column statistics are indexed; raw values are not, deliberately — which is what makes the catalog safe to show everybody.
3.9 External Assets
Role: every role can read the register; registering, editing and retiring an external asset requires Steward, Data Owner, Developer or Admin.
Your estate does not stop at the platform's edge. A finance warehouse consumes your Party golden records; a downstream report is built on a table nobody here owns; a lake file feeds a system that feeds you. Modelling Studio → External Assets is the register of those things — the assets that matter to your lineage but live outside MDM Studio.
Nothing crawls. This is a curated register, not a discovery agent. The platform never reaches into your network to find out what exists; a steward records an asset because it matters, and that record is the assertion. The register holds what somebody decided to say, which is why it can be trusted as documentation.
How external assets are addressed. They use the catalog's URN scheme with their own first segment: ac:x/{system} for a system, and ac:x/{system}/{name} for a named asset inside it. So ac:x/FinanceDW is the warehouse, and ac:x/FinanceDW/DIM_CUSTOMER is a table in it. That string is the join key everywhere else — in the catalog, in glossary bindings, in mapping specifications and in exports.
They are global, on purpose. An external asset carries no model. A finance warehouse is not a fact about one version of your design, and duplicating it per model would be three copies of one truth waiting to disagree. It is registered once and visible from every model. (The consequence to know: a colleague working in another model sees the same register, and a retire is felt everywhere.)
What you can record
- The kind of asset — system, table, file, report, dashboard, API and so on. The list on the form is not typed out anywhere: it is derived from the definition that decides it, so the picker cannot drift away from what the platform will accept.
- Description, owner and system, and a free set of the same curated facets the catalog carries.
- Links to what you already have — three authored kinds, and the direction is the meaning:
- CONSUMES — the external asset reads from your model (
ac:x/FinanceDW/DIM_CUSTOMERconsumesac:entity/PARTY). - FEEDS — the external asset supplies something of yours.
- DESCRIBES — the external asset documents or reports on it.
- A glossary binding. A business term can govern an external asset exactly as it governs one of your own columns, so Customer Dimension means the same thing on both sides of the boundary.
Registering is audited, not reviewed. Writing to the register is recorded in the audit trail with who and when, but it does not raise a change request. That is a deliberate asymmetry: recording that a warehouse exists is documentation, and putting a review queue in front of documentation is how registers go stale. Editing a catalog facet remains governed (§3.8) — a statement about ownership and meaning is reviewable; a statement that a table exists is not.
Retiring, and what a tombstone means. An external asset is retired, never quietly deleted. Retire states how many links it carries before you confirm, and it tells you the consequence: the asset stops being emitted into the catalog, so the next rebuild tombstones it there — marked as gone, with its history and its human‑set facets intact, rather than vanishing as if it had never been recorded.
The catalog is still a read model. External assets appear in catalog search, coverage and impact analysis as EXTERNAL‑type assets, and — this is the part worth understanding — the register is the authored home and reindex is the only door. Only ACTIVE assets are emitted, which is exactly the mechanism that produces the tombstone above. A row written straight into the catalog by some other route does not survive: the next rebuild tombstones it, because the catalog only agrees to hold what a module authored somewhere else first (§3.8).
Importing from an external catalog — proposals, and OFF by default. Where Catalog Integration (§10.18) pulls metadata in from Atlas, Purview or Collibra, that metadata can be offered to this register as steward proposals: register this as an external asset. Two things about that bridge:
- It proposes; it never registers. Every suggestion waits for an explicit accept, and a rejection is remembered.
- It ships switched off. A deployment that pulls from an external catalog does not start proposing external assets until an administrator turns the bridge on. A capability that was made possible is not a capability that was switched on.
Where the boundary is honest about itself. External assets are the only assets in the catalog whose truth the platform cannot check. Everything else in the index is projected from a module that owns it, so a wrong entry is a bug. An external asset is a human's statement about somewhere else, and it stays true only as long as somebody maintains it. Treat the register as documentation with an owner, and give it one.
3.10 Value Hierarchies
Role: every role can read a hierarchy and use it to navigate; declaring, binding and accepting containment requires Steward, Data Owner, Developer or Admin.
The third meaning of hierarchy (§3.7) is containment between values: South Africa contains Gauteng, Gauteng contains Johannesburg. Tiers cannot say this — a tier declares that PROVINCE is level 1 and CITY is level 2, but nothing in a tier records which province contains which city. A value hierarchy does, and that one fact is what makes a roll‑up trustworthy: a Johannesburg record under Western Cape stops being an unnoticed grouping quirk and becomes a named contradiction.
The shape. A value hierarchy is a named chain — GEO_ZA: Country ▸ Province ▸ City — in which every level is a governed reference domain (§4.3), and every containment link joins a value in one code set to its parent in the code set exactly one level above. Because the levels are governed code sets, the links are governed too: a link to a value that is not in the parent's code set is refused ("the link would dangle"), a cross‑level link is refused with the levels named, and a code set can never contain itself. Depth is declared, with a stated cap that refuses rather than truncating.
Declaring and binding are two separate acts. The Modelling Studio → Value Hierarchies page manages the chain (its levels, its links, its advisor); binding it to an entity — saying that Customer's PROVINCE and CITY attributes carry these levels — is per entity, and is deliberately a different binding from tier bindings: an entity can carry both, and a surface that filters by both AND‑combines two named filters, never one blended answer.
The advisor measures; it never guesses. Point the advisor at two bound levels and it reads your published golden records and proposes containment links with the evidence — how many records agree, how many disagree, how many are blank. A unanimous value is proposed; a contested value (Gqeberha claimed by two provinces) is proposed as nothing — it sits in the tree as a finding until a human resolves it. That restraint is the point: the two contested values are precisely the ones a guess would get wrong.
Two switches, both off until you choose them.
- Derive — display only. A record whose province is blank but whose city resolves under Gauteng can be shown under Gauteng, labelled as derived. Nothing is written; the toggle is off until you turn it on, and every surface that includes derived records says so.
- The
IN_HIERARCHYrule operator — a DQ rule (§6.4) can flag records whose two bound values contradict the declared containment. It flags contradiction only: a blank parent is a completeness question and an unknown value is a mapping question, and each has its own finding. No rule ships using the operator — you author one when you want the check. A rule whose hierarchy is missing or capped fails its run rather than passing every record.
Where you meet value hierarchies. On the ERD, as containment edges and chips (§3.1). In the GR Explorer's hierarchy rail, as the Value hierarchy source (§8.4) — picking Gauteng scopes the grid to the branch. On the Record Inspector, as a per‑record path line (South Africa ▸ Gauteng ▸ Pretoria), which is how the containment reaches every golden surface and the Exception Manager. And in Data Quality's By‑hierarchy view, wherever the entity has a navigable source.
4. Integration Hub
The Integration Hub connects your sources and moves their data into the model, stage by stage: Extract → Land → Stage → Standardize → Match → Survive → Publish.
4.1 Sources & Connections
The Sources & Connections tab has three parts:
- Source Systems — logical systems of record (CRM, ERP, Billing, HR …).
- Connections — the credentialed connections used to read them.
- Source Objects — the registered inventory of the tables, files and endpoints you actually read (below).
MDM Studio supports fifteen connection types across four families:
- Relational databases — SQL Server, Azure SQL, Azure Synapse, PostgreSQL, MySQL, Oracle, Teradata.
- Cloud data platforms — Snowflake, BigQuery, Databricks, and Salesforce as a SaaS source.
- Files — CSV and Excel (with a worksheet chooser).
- Integration — FTP file drop/pickup and REST APIs.
The five cloud/SaaS connectors ship as optional components. Snowflake, BigQuery, Databricks, Salesforce and Teradata each need a driver that a standard installation does not include. If one of those types is missing from the picker, or a connection to it fails immediately with a driver error, that is the answer — ask your administrator to install the optional driver for that platform.
Connections live in one global registry shared across every hub engine — created once, replicated everywhere, and self‑healing (a connection defined on one engine simply works on another). Credentials are encrypted at rest, and each source can carry a trust score used later during survivorship.
The AI can draft a profile for a source system or connection — a suggested description and a defensible default trust score based on the system type (ERP, CRM, HR, files …), with its rationale shown; file and feed types cap the suggested trust. Duplicating a connection asks for confirmation, and connection exports include a sanitised connection string per engine — credentials are never exported.
Source Objects. The third tab registers the source objects — the specific tables and files your pipelines read — and snapshots each one's field inventory. Mark fields in or out of scope (coverage KPIs always show their denominators), and re‑snapshot at any time to detect drift: added, changed and removed fields are listed explicitly, removed fields are kept struck‑through rather than silently dropped, and a removed field that a live mapping still reads is flagged BREAKING by name. This is where "the source changed under us" stops being a surprise.
Role: creating/editing connections generally requires Developer, Steward, Data Owner or Admin.
4.2 Mapping Designer
The Mapping Designer maps source fields to your model's attributes, stage by stage through the pipeline. Preview source schemas, sheets and endpoints with live previews, then map fields into your model by dragging source columns onto attributes.
Stage rules (enforced). An Extract set reads from a registered source connection; every later stage (Land, Stage, Standardize, Match, Survive, Publish) reads only the previous stage's output. The create‑set form makes this concrete: for Extract it introspects your source live and offers schema/table dropdowns; for later stages it offers a single previous stage output choice and prefills the name and description. Invalid combinations are rejected with instructive messages rather than failing downstream.
Build the whole pipeline in one click. Once your Extract sets exist, Build pipeline sets generates the full stage‑to‑stage lineage for each entity — idempotently, healing crossed or stale sets and warning about entity‑name collisions. Deduce (per stage tab) creates sets in the stage you're looking at.
Loading behaviour. Every Extract set carries a load‑expectation badge: Delta BAU (incremental on a chosen date column — the form offers the candidate date columns), Full‑load BAU (acknowledged), or an amber undecided nudge you can acknowledge in one click. A clear amber warning also appears when a set has mappings but no key field flagged, explaining the consequences for matching and golden records.
Multi‑source tables extract into per‑entity tables, so two entities fed by the same source table never contaminate each other's pipelines.
Transforms execute — and the designer is honest about the two that do not. A per‑field transform (UPPER, LOWER, TRIM, an EXPRESSION) is applied by the pipeline's generated extract SQL, so what you configure is what lands. Two transforms — CONCAT and LOOKUP — do not execute at run time, and the designer says so: their picker entries are labelled — not applied with the reason, saving a mapping that carries a non‑executing or unknown transform is refused with a suggested alternative, and existing saved rows that carry one show an amber warning. A transform‑impact report lists every existing mapping whose output changes when transforms run — read it before upgrading a hub that predates transform execution (§12.7).
Specification versions — what actually runs. Approving a mapping set freezes an immutable, numbered specification version, and the pipeline executes the approved version, not the editable draft. The draft stays editable at all times; the Versions panel states which of three states the set is in — no approved version (the draft runs, exactly as before), approved and matching (draft and approved agree), or diverged (the approved version still runs until you approve again) — with a version list and a field‑level diff showing both values for every change. Approving an unchanged draft is refused; each version records who froze it and when. A spec‑impact report names the sets whose drafts have moved on since approval (§12.7).
A specification can target something outside the platform. A mapping set names a target kind — ENTITY (one of your model's entities, the normal case) or EXTERNAL (a registered external asset, §3.9). An EXTERNAL‑target specification is how you write down, and govern, the mapping from your hub to a downstream warehouse table: Hub Party → FinanceDW DIM_CUSTOMER, field by field. The Designer's create form carries the target‑kind picker, and the two are mutually exclusive by construction — a set targets an entity or an external asset, never both and never neither.
These specifications are documentation with teeth, not a second pipeline:
- They are global, like the external assets they point at (§3.9) — not scoped to one model.
- They are governed and spec‑versioned exactly like an entity‑target set: approval freezes an immutable numbered version, with the same diff and the same diverged state.
- No pipeline will run one. Ingestion refuses to create a job for an EXTERNAL‑target set, the pipeline builder refuses to generate stages for it, and the field guard refuses it in both places — each by name, saying which set and why. The platform will not quietly do half of something it cannot finish; MDM Studio does not write to a warehouse because a mapping set described one. Moving data outward is Write‑Back (§5.3), which is a separate, deliberately chosen mechanism.
They also appear in the catalog as MAPPING_SET assets and contribute external FEEDS edges, so an impact analysis can follow your model outward to the systems that read it.
The Scripts tab previews; it does not promise. Every SQL box on the Scripts tab is read‑only: the Extract tab shows the same builder the pipeline runs, and the DDL and MERGE tabs are labelled Preview only. There is no Save — saved custom SQL was never executed by the pipeline, and the tab no longer implies otherwise.
Mapped / Unmapped are stage‑accurate. An attribute counts as mapped only if it has sources in the selected stage, and post‑Extract stages show the mapping feed from the previous stage's schema — Standardize reads the Stage output, Match reads Standardize, and so on. Where an attribute has active standardization rules, MATCH‑stage proposals automatically read the standardized STD_ columns, and revert to the raw columns if the cascade is deactivated. Hand‑crafted source bindings are left alone, and proposals say how many they left as they are.
4.3 Standardization
The Standardization tab defines rules that normalise incoming values — casing, formats, string operations and lookups against governed reference data — so that matching and quality work on clean, consistent data. You can preview and test standardisation cascades before applying them. Reference lookups support whole‑word and substring replace semantics (explained inline), and a lookup whose reference domain is missing or inactive fails loudly rather than being silently skipped.
An execution strip on the page shows the last Standardize run (rows processed, rules applied, lookup misses) with a Run Standardize button and a link to the owning job. Saving a rule tells you how many entities' pipelines become stale — and stale stages are flagged on the pipeline board with one‑click re‑run (§5.4).
A Standardization Advisor reads your profiled data and proposes cleansing cascades attribute by attribute, with the evidence for each proposal shown. Like every advisor in the product bar the ones listed in §12.6, it is deterministic — the same data produces the same advice, and you can read why.
The rule builder. Creating or editing a rule opens a builder that grounds every suggestion in something you can check. Candidates arrive in layers, each naming its provenance: the governed layer (the attribute's logical type and reference domain, read live from the catalog — never a copy), the evidence layer (every DQ rule on the column, with its fix pairing — see §6.4), and the pattern layer (the curated transform library). A suggestion that cannot be attributed cannot be judged, so every candidate carries its layer and source on its face.
Describe what the rule should do. The builder starts with an intent box: type "remove dialling code e.g. +27 / 0027" and the platform reads the verbs and the literals out of the sentence and composes a portable expression — one that runs identically on all four hub engines. This layer is deterministic and needs no AI key: it knows remove, replace, digits‑only, trim, casing, pad, truncate and null‑if‑blank, and what it cannot do portably it says rather than drops. Its candidates come first, badged DERIVED — this came from your sentence, not from anything anybody approved — and the box states what it read back to you (Read from your words: remove '0027', '+27'). Where a sentence has two honest readings — delete +27, or substitute the national trunk prefix — both are offered and the choice is named, because a confident single answer would simply be one of them, half the time wrongly. With an AI key configured, a language model adds candidates on top; §12.6's rules on what is and is not sent apply.
Try one value. Type a single value and every candidate on screen — including whatever is in the expression box — shows your value → result, computed through the same compiler and executor the pipeline runs, so what you see is what STANDARDIZE would produce. Three outcomes, all named: transformed, unchanged (a real and common answer, said in a word rather than leaving you to compare strings character by character), and failed with the database's own words. Nothing is written; no population is read.
Preview against the real column before you commit. A transform can fix every failing value and still be wrong — a left‑pad that turns '99' into '0099' also turns '12345' into '2345', which passes the length check and is a different number. The preview runs the whole cascade over the column's real population and buckets every value: fixed, still failing, broken (was passing, now fails) and rewritten (still passes, but the stored value changed — the quiet one). It answers with a verdict sentence — "Fixes 805 failing record(s) and CHANGES 421 record(s) that were already passing" — never a coverage percentage, and it refuses on a masked column. A transform the product proposed cannot be saved until it has been previewed against the real column, and editing the expression afterwards invalidates the preview; an expression you typed yourself saves as it always has.
Reference data lives here too. Standardization carries a Reference Data tab and a Managed Values tab; these are the one destination for governed reference domains and managed values, and the former standalone pages redirect here. Reference domains, managed values and standardization rules all show "used by" impact chips with navigable usage panels; deletes are guarded by a real usage breakdown, and retiring a reference domain requires an explicit, audited confirmation.
Deleting a standardization rule does not delete it immediately. The rule is soft‑deleted and a change request is raised for the removal, so a rule that other work depends on cannot vanish without review.
4.4 Jobs & Runs
The Jobs & Runs tab is the pipeline's operations console: define and run the jobs that move data stage to stage, and see every run in one place.
Creating jobs. The quickest way is to drag mapping sets from the palette onto the jobs panel — drag one, or tick several and drag them together; each drop creates the job for that set silently, and wrong‑stage or duplicate drops are blocked per set with the reason shown. (The + New job button keeps the full form for fine control: source connection, mapping set, target entity, load strategy, schedule.) Auto‑allocate goes further: one click proposes moving existing jobs to their logical pipeline stage and creating jobs for mapping sets that have none — nothing changes until you confirm the previewed plan.
Runs and feedback. Every run stores a structured summary. Success dialogs and run history flag warnings — rows stamped WARN, soft‑retired rows, capped reads — and run logs are colour‑coded. The All Runs view unifies job, rule, match and profiling runs in one filterable history (by kind, entity and status).
Record lifecycle. Stage rows carry a status — STAGED, WARN or RETIRED — visible as badges and filters in stage reports and as chips on the pipeline board. Retired rows never flow to later stages; when a record disappears from its source its golden record is superseded, and it revives automatically if the record returns (§5.2).
4.4.1 Run notifications
You can subscribe to a job, a job group or an execution package and be told when it starts, succeeds or fails — in the app, by email, or both. Subscriptions are per person; the buttons sit on the job and group pages beside the run controls.
One incident sends one message. Where you are subscribed at more than one level — watching a job and also the group that contains it, which is the normal thing to do — the most specific scope wins. The job names the thing that actually failed, so the group and package stay quiet for that run rather than telling you the same news three times.
4.5 Orchestration
The Orchestration tab groups jobs for coordinated execution:
- Job Groups — run a set of related jobs together. Multi‑select jobs and drag them in; every drop opens an Add to group prompt (pick an existing group or create one inline). Group cards are collapsible, and an auto‑allocate planner can build groups for you — by domain ("Domain pipeline", jobs in stage order) or by stage — with a live preview of the plan before anything is created.
- Execution Packages — package a sequence of steps as a repeatable unit. The palette works the same way (multi‑select groups or jobs, drop, pick a target package), package entries expand in place to peek at their contents, and the Standard pipeline button composes an entity's existing stage jobs into one canonical, dependency‑chained package — reporting any stages that are missing.
Groups and packages roll up run feedback ("N jobs with warnings") so a quiet green run really means quiet and green. Every job in a package runs under the package's run group, and All Runs shows the parent pipeline's start time beside each member, so a package's steps read as one run rather than a scatter of independent ones. A package that cannot prove a job belongs to your tenant stops with a FAILED row and a reason, never a silent hang.
5. Mastering
Mastering turns many source records into trusted golden records.
5.1 Mastering Rules
The Mastering Rules tab holds the three ingredients of a golden record:
- Matching Rules — rulesets that score whether two records represent the same real‑world thing. Each rule compares attributes with a chosen comparison method, weight and threshold; you can use blocking keys to keep comparisons efficient, and review candidate pairs. The comparators are Exact, Normalized, Prefix, Fuzzy (edit distance), Damerau‑Levenshtein (which also forgives transposed characters), Jaro‑Winkler (threshold‑based, favours a shared beginning), Metaphone (phonetic) and Date tolerance (dates within N days). A scored pair resolves to one of three decisions: AUTO_MERGE, REVIEW or NO_MATCH — only the middle one asks for a human, and it arrives in My Work as a match review.
- Survivorship — per‑attribute and per‑entity strategies that decide which source value wins: most trusted source, most recent, most complete, longest value, most frequent, and CUSTOM. The CUSTOM strategy evaluates real expressions with
@Columnplaceholders (aggregate and scalar forms); the strategy picker is level‑aware (CUSTOM per attribute, MOST_COMPLETE at entity level), and a warning fires if Most trusted runs while no source trust scores are configured. Where sources disagree in a way no strategy resolves, the conflict is recorded and raised as a survivorship conflict in My Work. - Golden Fields — which attributes are included in the published golden record (driven by golden‑record scoring in the Data Dictionary), and whether each attribute masters its standardized value (below).
A Match Advisor and a Survivorship Advisor propose rules from your profiled data and explain each proposal. Both are deterministic (§12.6): they read your data, not a model, and they never apply anything on their own.
For large volumes, a ruleset can enable set‑based candidate generation: the hub database pre‑filters the rows that could possibly match before any pairwise scoring, producing identical results with far less load — rows with no candidate partner are marked unique without ever being loaded.
Mastering the standardized value. An attribute with standardization rules produces a cleansed value — but the golden record only carries it when the attribute's master standardized value flag is on. Off, the rules still run and matching still uses the cleansed form, yet every published golden record keeps the raw survivor: the cleansing is invisible to every consumer. Three things keep that from happening quietly:
- The default. The first time a standardization rule is attached to an eligible attribute on which nobody has yet decided, the flag defaults on — so a new rule's cleansed value reaches the golden record without a second step. The decision is recorded (by whom, when), and the three states stay distinguishable forever: undecided, on, and off by somebody's explicit choice. Every surface that shows an attribute with standardization — the Data Dictionary, the tiers page, the DQ lens — carries a small mastering indicator that renders exactly that verdict (ON, OFF, undecided, or interim — standardized for grouping while the stored value is still raw). The badge shows the server's answer; it computes nothing of its own, so two screens can never disagree about it. A recorded human "off" is never overturned by the default, no matter how many rules are attached later.
- The offer. Attributes that predate the default are offered, never flipped — switching a column on rewrites its published values at the next Publish, which is a decision, not a migration. The offer panel (on Golden Fields, and on the tiers page for tier‑supplying attributes, §3.7) lists the candidates; each is previewed against its real data before applying, and a column whose cleansed value would change records that currently pass their rules is refused until you acknowledge that collateral by name. Where no data‑quality rule judges the column, the collateral is reported as unknown — which is not zero — and also needs an explicit acknowledgement.
- The disclosure. Until a tier‑supplying attribute is mastered, its hierarchy tier groups by the cleansed value while the stored column keeps the raw survivor (§3.7); the surfaces that show tiers say so, and the offer is the way to make them agree.
Switching mastering on changes nothing by itself — the cleansed value is published at the next Survive + Publish for that entity.
An input status strip on the matching page shows whether MATCH‑stage data has been built; if a run fails for lack of it, the message explains that loading the MATCH stage is an ingestion job, with direct links to fix it.
Runs refuse rather than truncate. Survive and publish runs carry row caps. Past the cap the run fails with a clear message; it never quietly processes part of your data and reports success. If you hit a cap, ask your administrator to raise it for that deployment rather than re‑running and hoping.
Role: Steward, Data Owner, Developer or Admin.
5.2 Golden Records
The Golden Records browser shows the mastered output. For each record you can:
- See the golden value for every field, which source it was won from, and the survivorship strategy that chose it. When the golden value differs from the raw source value, the raw value is shown alongside ("was: ZA") — in the app and in CSV exports — and per‑attribute chips show the standardization cascade that transformed it.
- See the record's confidence — computed from real match evidence (pair scores and field agreement; singletons are fixed at 50) — with a full breakdown in the tooltip, and open View match evidence for the pairs behind the cluster.
- Expand the contributing source records (the cross‑reference/xref lineage) behind the golden record, or open the Record Inspector overlay for the full picture — cluster members, related records, exceptions and lineage in one place. A record that matched nothing says so plainly: it forms a cluster of one — no duplicates found.
- Merge two golden records or split one, and view history.
- Restrict a record so only cleared users can see it (Chapter 11) — useful for VIPs, employees or minors. Restriction works before publishing too: restricting a not‑yet‑published record creates a hidden stub that keeps its restriction through survive and publish cycles.
- Jump to the record's open exceptions, or Analyze in Report Studio — a new report preloaded with the entity's golden data source.
Record status. Publishing never deletes: when a source record disappears, its golden record is marked SUPERSEDED (and revived automatically if the record returns). The browser filters by status — Active, Superseded, Merged, Unpublished (restricted stubs) — with badges and a clickable superseded count. Golden keys are shown trimmed (~10 characters, hover for the full key); singleton records use the short stable format U-<system>-<hex>.
Scope the browser by hierarchy. Entities with a navigable hierarchy gain a Hierarchy control here too — the same rail as the GR Explorer (§8.4), with the same sources and the same behaviour. One nuance is specific to this page: because it lists the pre‑publish survived rows, a rail pick is evaluated against the published golden records — a survived row is in the branch when its published golden record is — so the branch means the same thing here as everywhere else. Survived rows not yet published are outside any branch while a scope is active, and the filter chip says so. An oversized branch blocks the query with its reason rather than quietly showing the whole entity under the branch's name.
Values above your sensitivity clearance appear masked, and restricted records you're not cleared for are hidden entirely.
5.3 Write‑Back
The Write‑Back tab returns mastered values to the source systems they came from, under governed control.
How a write‑back batch runs. First you make attributes eligible — opt in per attribute or per entity, and name the role with authority over it. Then: build a batch of proposed changes, preview them, run an impact check, approve or reject each item, apply, and verify. If verification is not what you expected, there is a rollback, because before‑images are kept.
Read this before you promise anything to a source‑system owner.
- The default is a script, not a write. Out of the box a batch produces an idempotent, drift‑guarded UPDATE script for a source‑system DBA to run — the platform changes nothing itself. The script re‑checks that the row still looks as it did before it updates, so it is safe to hand over and safe to run late.
- Two other modes exist and must be chosen deliberately: PROC (call a stored procedure the source team owns) and DIRECT (the platform executes guarded updates itself, in chunked transactions, keeping before‑images for rollback).
- Write‑back is not available on an Oracle hub. If your master‑data hub runs on Oracle the feature returns a clear "not available on this engine" message. It is supported on SQL Server, PostgreSQL and MySQL hubs. (This is about the hub engine, not the source you are writing to.)
Write‑back is wired into the stewardship loop: an exception whose failing value has a known fix can offer Remediate via write‑back, prefilling a batch; each write‑back item shows which exception it resolves, and successful verification auto‑resolves that exception (§6.3).
5.4 Pipeline / Lineage
The Pipeline / Lineage tab is the per‑entity pipeline board: every stage of the flow for an entity — ext → LND → STG → STD → MTC → SRV → GRD — with row counts, run states and the movement of records between stages, plus end‑to‑end provenance for a golden record and its fields.
- A hoverable stage glossary explains each stage code.
- Stages whose configuration changed since their last run are flagged "Stale — config changed" with a one‑click re‑run.
- The record‑flow chart overlays an actual‑counts line on the stage‑movement bars, so expected and actual movement can be compared at a glance.
- Deep links can preselect an entity and stage — other pages (profiling, matching, run feedback) link straight into the board.
5.5 Author Records
Role: Admin, Steward or Data Owner.
Not every master record has a source system. Mastering → Author Records lets a steward create or amend a golden record directly in the hub — the new subsidiary that exists in law before it exists in the ERP, the reference customer nobody's system owns.
- The form is built from the entity itself: its attributes, their types, their reference domains. Nothing has to be configured for a new entity to be authorable.
- Values are validated against that entity's live DQ rules as you enter them, so an authored record cannot be created already in breach of the rules everything else is held to.
- The record is stamped with MANUAL provenance, which is visible everywhere lineage is shown — an authored value is never mistaken for a value that survived from a source.
- If an approval workflow covers the entity, the record is queued for approval rather than published (§7.6), and the approver sees it in My Work.
Authored records live in the same golden set as mastered ones: they match, they can be merged, they carry exceptions, and they appear in GR Explorer and reports.
5.6 Record Timeline
Role: any role that can view the entity's golden records; masking applies as usual.
Mastering → Record Timeline answers questions about time, and it distinguishes two kinds of time that are easy to confuse:
- As‑of — what was true in the world on this date? The customer's address as it was in March.
- As‑known‑at — what did the platform believe on this date? The address we would have printed on a March invoice, even if we have since learned it was wrong.
Both are held, which is why the timeline can answer "we sent that letter to the wrong address — was it wrong, or did we not know yet?" A record's timeline shows every version with its effective dates and the change that produced it, and you can pin either axis to reconstruct the record as any point in the past.
5.7 Real‑time API
Role: Admin, Developer or Steward to configure; the API itself is called by systems, not people.
Mastering → Real‑time API exposes two synchronous operations to your other applications:
| Operation | What it does |
|---|---|
| Match | Takes an inbound record, scores it against the published golden master using that entity's active matching ruleset, and returns ranked candidates with a decision of AUTO_MERGE, REVIEW or NO_MATCH. |
| Golden lookup | Returns the published golden record for an entity and key. |
Both use the same scoring engine as the batch matcher, so a real‑time decision and a batch decision on the same pair agree.
Be precise about what this is. The real‑time API decides and serves; it does not master. A real‑time match tells a calling system "this looks like customer 4471, with a score of 0.91" — it does not create, merge or update a golden record. Mastering is the batch pipeline: Match → Survive → Publish, run on a schedule, by hand, or as part of a job group. What the API serves is the last published master. If a source changed an hour ago and the pipeline has not run since, the API will faithfully serve what was published before that change.
Two further limits worth knowing:
- A synchronous match loads at most 5,000 golden candidate rows. Blocking keys are what keep a real entity inside that bound; without them, a large entity will hit it.
- Event intake is a queue, not a stream. Posting an event durably enqueues it; processing is an explicit call (one event or the whole queue). That makes it replayable and auditable — it does not make it a stream processor.
Records the API decides need a human arrive in My Work as real‑time reviews.
For machine consumers there is a separate token‑authenticated API surface, so an application integrates with a service token rather than a user's session. Administrators issue and revoke those tokens.
6. Quality
The Quality hub measures and improves the trustworthiness of your data.
6.1 Data Profiling
The Data Profiling tab scans a stage table and reports column‑level statistics: population and null rates, distinct counts, value patterns, lengths, min/max and a value‑frequency distribution. Profiling helps you spot data issues before (or after) mastering — its place in the flow is right after ingestion: ingest → profile → standardize.
Profiling detects the furthest‑built stage automatically (including Extract‑only pipelines) on every engine, excludes retired rows (and reports how many were excluded), and its error messages name exactly what was checked. (For sensitive columns above your clearance, the value list and min/max are suppressed while the statistics remain.)
6.2 Data Quality
The Data Quality tab is the scorecard: validation rules and rulesets produce quality scores per domain, entity, hierarchy tier and dimension, captured as snapshots over time so you can track trends and SLAs. Scores refresh automatically after rule runs and publishes, with an "as of" timestamp so you always know how current the number is.
The report is organised as tabs — Overview, Entities, By tier, By dimension, By hierarchy, Coverage, Targets and Domain Health — and two behaviours run through all of them:
- Click to filter. Clicking a domain or entity in the scope tree selects and filters — the grid, the KPI tiles and the trend all follow the narrowest selection, and one scope row states what you are looking at with a clear control to clear it. Nothing silently keeps showing the whole estate while the header claims a scope.
- Dimensions are named, not just coloured. Every score strip carries the dimension's letter (C·V·U·Cn·I·A·T) with a legend on the page, so you can tell completeness from consistency without hovering a colour block. Grey means not measured — which is not a pass and not a zero.
By tier groups golden records entity → tier level → values: an entity header with its record‑weighted rollup, level sub‑groups ("Tier 1 · 2 values"), and one scored row per value. Record‑count pills show how many records carry a value at each node; a tier row whose level was not recorded lands in a named "(level unknown — recompute scores)" group rather than being guessed into a level — one recompute stamps it.
By dimension turns the report sideways: one row per quality dimension across the scope, so "where is consistency worst" is one glance rather than a mental transpose.
By hierarchy scores quality down a tree, and appears once the entity has any navigable hierarchy source. It offers the same three lenses as GR Explorer (§8.4) — Managed tree, Record tree and Tiers — never blended, each disabled with its reason written on it when that source does not exist for the entity. Tree nodes carry their own score, their own record pill and a "N below" subtree pill; the node labels follow the hierarchy's label attribute (§3.7), editable right on this page. The Tiers lens drills level by level in place — expanding a value loads the next level's scored values beneath it, sibling branches stay open side by side, and a "(no value)" row is a finding, not something you can expand.
6.3 Exceptions
The Exceptions tab is the remediation workbench. Each rule failure becomes an exception with a lifecycle:
OPEN → ASSIGNED → IN PROGRESS → RESOLVED (or IGNORED), with reopen and SLA‑breach states.
New exceptions are auto‑assigned to the responsible steward (attribute → entity → domain) with an SLA due date by severity. You can filter by status, severity, domain, entity, rule, assignee and more; assign, start, resolve or ignore; and trace each failing value back to its contributing source records. Role: Admin, Steward, Data Owner or Analyst.
The rollup tree. The left of the page is a tree of every open exception grouped domain → entity, and each entity opens into its quality dimensions and severities. It is a triage instrument: you can see at a glance that the 4,000 open exceptions are 3,700 completeness failures on one entity, and go straight there. Selecting a node filters the grid, and the same selection can be reached by link — so a dashboard tile, a report or a colleague's message can hand you exactly the slice they were looking at.
Scope the list by hierarchy. Once the tree has narrowed to one entity, a Hierarchy control offers the same rail as the GR Explorer (§8.4): scope the exception list to a tier value, a branch of the record tree, or a value‑hierarchy branch. An exception is in a branch when its golden record is — so exceptions with no golden link (a source‑record finding, or a record whose golden was re‑published away) are excluded while a scope is active, and the filter chip's footnote says so. A branch too large to resolve blocks the query with its reason rather than silently widening to the whole list.
Group by rule. A queue often repeats one rule name down hundreds of rows. The Group by rule toggle collapses them into one row per rule, with the rows underneath opening on click. It is a view, not a filter: it arranges what the filters left and removes nothing, and its chip in the filter bar says so. Every group header reports two numbers, never one merged figure — open, which is what the rollup tree counts, and closed — plus its scope ("12,004 open in Sales ▸ Customer"). Select all becomes Select all shown and reads only the rows in expanded groups; collapsing a group deselects its rows, so a bulk action can never touch a row you cannot see. Open this rule on its own sets the ordinary rule filter.
Working in bulk. Assign or resolve many exceptions at once. Bulk transitions run in the database rather than record by record, so a large sweep is quick and has no row limit. One case is deliberately excluded: reopening depends on a record's human resolution history, so those transitions are handled individually and reported back separately as deferred — a bulk action tells you plainly which items it did not decide.
Resolution quality. The page also reports bounce rate — how often a steward's resolutions came straight back. It is a measure of whether remediation is working, not a measure of the steward; a high bounce rate on one rule usually means the rule or the standardisation behind it is wrong.
Typed resolutions with re‑validation. Resolving asks how it was fixed — Corrected, Fixed at source or Accepted — and then re‑runs just that rule on just that record. If the value still fails, the exception reopens immediately with an ineffective resolution event; Accepted skips the recheck by design.
Auto‑remediation. A rule flagged auto‑remediate applies the attribute's standardization cascade server‑side when a violation is found — correcting the golden record and closing the exception as Corrected, no human touch needed.
A fully linked loop. A golden record's detail lists its exceptions; an exception's detail links to the golden record, the rule that fired, and its campaign; per‑rule open‑exception counts drill through to the filtered list; and exceptions carrying a fixable value offer Remediate via write‑back (§5.3).
6.4 DQ Rules
The DQ Rules tab authors the validation rules that generate exceptions — checks over attributes, with dimensions (validity, completeness, consistency, …), severities and pass thresholds. Every rule gets a unique, auto‑numbered code (DQ-#######_<your code>), so duplicate codes are impossible. Rules created before auto‑numbering keep working under their old codes — saving one stamps it into the convention, and a Renumber legacy codes button appears on the toolbar whenever any remain, converting them all in one audited step.
See the code behind a rule. In the rule editor, View rule code shows exactly what a declarative rule stores (operator, parameters, threshold) and the equivalent SQL for its failing‑rows query — useful for review, audit and hand‑off to database teams. Blank values never fail content checks (completeness is authored as its own rule), and the SQL view says so.
Authoring made easy.
- AI suggest infers everything from the rule name and target attribute: the DQ dimension, the check operator and parameters, an error message, a threshold and a rule‑set match. It understands plain language — "Should not contain commas or dots" becomes the right pattern (spaces still allowed, blanks failing by default) — backed by a curated, self‑tested regex pattern library (emails, phone numbers, national IDs, IBAN/SWIFT, postal codes, casing checks and many more).
- Declarative operators cover the common cases with no regex at all: starts/ends with, contains (and negations, each with a case toggle), casing checks (UPPER / lower / Title / Sentence) and trimmed‑whitespace.
- An (i) panel explains each DQ dimension with example violations, and the Dimensions guide lists all seven side by side.
- A distinct‑value preview beside the target attribute shows real values from the furthest‑built pipeline stage (labelled, e.g. from Published golden records), so you can see what you're validating against.
Deriving the fix from the failure. A DQ rule and a standardization rule are near‑opposites — one detects a bad value, the other rewrites it — and the exceptions a rule raises are, by construction, the exact specification of what the cascade did not cover. The Fix action (the wrench) on a targeted rule turns that into a governed workflow. It reads the rule's real failing values (distinct, with counts, the cap stated), pairs the rule's operator with a candidate transform from a governed registry, and proves the candidate on the bench: the transform runs over the real failing values and the rule's own predicate re‑evaluates each output, so the answer is "this fixed 805 of the 805 exceptions these values account for" — with its denominator, never a bare percentage. Three pairing kinds, each honest about itself: FIXES (the transform makes failing values pass — refused if it proved nothing), ENABLES (it makes the data agree with what the check found — nulling a whitespace‑only value fixes zero exceptions, and the panel says that zero is the correct outcome), and NONE (there is no safe transform, and the entry exists to say why — seventeen operators answer this way rather than inventing something). Accepting creates the standardization rule inactive, assigned, and routed through the same governance a hand‑authored rule takes; the server re‑derives and re‑proves before creating, so a caller cannot assert its own gate. And because a bench that reads only failing values is blind to the records that were fine, the accept is wired to the preview (§4.3): a fix with collateral — passing records it would rewrite — is refused until you acknowledge that collateral by name.
Organising rules. The Rules tab is grouped by default — collapsible domain › entity › dimension bands, with a Flat toggle for the classic table. On the Rule Sets tab, the rules browser shows every rule in a collapsible tree with a seven‑way grouping selector; multi‑select rules and drag them onto an existing set — or drop them on the create zone to start a new set with an inferred name.
6.5 Campaigns
The Campaigns tab organises remediation into managed efforts — group exceptions into a campaign, assign owners and due dates, and track progress to completion.
Campaign progress is honest: it splits resolved / ignored / bounced, and ignored items don't count as progress. A campaign snapshots its pass rate at launch and reports "at launch X% → now Y%", so the improvement is measurable; completing a campaign with open items warns first. Work this campaign opens a queue that walks you through the open exceptions one by one.
The platform also suggests campaigns from your exception history, grouped by severity. Suggestions appear only once there is history to read — a new deployment has none, and that is expected.
A campaign has no delete. Campaigns are the record of a remediation effort, so they are completed or abandoned, not removed.
6.6 Rule suspension
Role: Steward, Data Owner, Developer or Admin. Suspending is treated as an operational action, so it does not require a model checkout.
Turning a noisy rule off with its Active switch loses two things: why it was turned off, and any prospect of it coming back. A suspension keeps both. On the DQ Rules page, suspend a rule with an end date and a reason — both are mandatory — and:
- The rule stays configured, visible and in its rule set. Nothing is deleted or unlinked.
- It returns by itself when the window ends, and the platform records that it resumed.
- A suspended rule is not a passing rule. This is the important part. It is reported exactly like a rule that could not be evaluated: no pass rate, no exceptions, no effect on the score. The aggregate states how many rules it did not evaluate, so a suspended rule can never quietly flatter a scorecard.
Four things happen without you asking:
- An expired suspension resumes, and says so.
- A rule whose target attribute has disappeared from the model is auto‑suspended rather than left to fail every run.
- A persistently noisy rule — one passing 1% or less across three consecutive runs — produces a proposal to suspend it. The panel says so in as many words: nothing is suspended by this. The scan is bounded, worst‑first, and reports its own cap, so on a very large estate it is telling you about the worst rules it examined, not necessarily every rule that qualifies.
- You are warned before a suspension expires, so a rule does not come back mid‑cycle by surprise.
The last two reach you as separate kinds of work in My Work, and the distinction is intentional: a suspension expiry is a deadline the system will honour on its own, while a noisy‑rule proposal is a question with no deadline at all.
6.7 Placeholder Advisor
Role: Steward, Data Owner, Developer or Admin.
Placeholder values — UNKNOWN, N/A, 0000000000, noreply@, a default date — are the quiet killer of a match ruleset. Ten thousand records sharing a placeholder identity number will match each other perfectly, and merge into a monster cluster.
Quality → Placeholder Advisor reads your profiling run and finds values that behave like placeholders in columns that are being used as match keys. For each one it shows the evidence: the value, how often it occurs, which column and which rule uses it. You then decide, per finding, whether to exclude it — nothing is applied until you review and accept it.
Because it reads a profile, it needs one: run Data Profiling (§6.1) on the entity first, or the advisor will have nothing to report.
6.8 Source DQ
Role: Steward, Data Owner, Developer or Admin.
Quality → Source DQ runs data‑quality checks directly against a raw source table, with no model at all. It is the tool for the conversation that happens before a programme is approved: point it at a table in a registered connection, run checks, and produce evidence of what the data actually looks like today.
Because there is no model, there are no entities, no golden records and no exceptions in the usual sense — the output is a report on the source. A check that proves useful here can be promoted into a real DQ rule (§6.4) once the entity exists, so the assessment work is not thrown away.
Source DQ and Source Comparison are the two screens in the product whose availability depends on your edition (§10.19). If they are missing, that is a licensing setting, not a fault.
6.9 Source Comparison
Role: Steward, Data Owner, Developer or Admin. Comparing raw source tables (§6.9.2) needs Admin or Developer.
Quality → Source Comparison answers "where do our systems disagree?". It lines up the records that two or more sources hold for the same real‑world thing and shows the fields where they differ — CRM says one address, billing says another. Every field is compared under a method you choose (exact, case‑insensitive, normalized, fuzzy, Jaro‑Winkler, numeric or date tolerance) and reported as AGREE, NORMALIZED (agree under the method), CONFLICT or MISSING.
It is read‑only: it reconciles and reports, it does not change either source. Differences that need a decision arrive in My Work as comparison conflicts, and the decision itself is made in survivorship (§5.1) or by write‑back (§5.3). A comparison set can be edited after it is created — its name, fields and methods — without losing its run history; its mode and its entity are fixed, because the runs and open discrepancies describe the comparison it does today.
6.9.1 Two modes, one artifact
Matched entity — the original mode. Comparison rides on the entity's existing match clusters: matching has already said which rows across sources are the same customer, so the comparison only has to diff the fields. It is scoped by the model like everything else, and it compares standardized values by default, so ZA and South Africa read as agreement rather than a false conflict.
Source tables — compare two or more raw tables straight off a connection, before any mapping or modelling. Pick a connection, a schema and the tables in the builder's navigator, give each an alias, and the set is built from a record key per table and a field map. A raw table is read outside the model — it inherits no domain and no row‑level access — which is why this mode needs the same role as creating a source connection, and why the builder says so where the tables are chosen. Both modes create the same comparison set and are read by the same list, run and results grid.
6.9.2 The record key — the one thing that cannot be omitted
A field map says which columns line up. It says nothing about which rows correspond — matched mode gets that from the match cluster, and table mode has to be told. So every table in a table‑mode set declares a record key: the column, or combination of columns, whose value identifies one row. Two rows are paired when their keys are equal, and the key must have the same number of parts on every table.
Derived keys. Two systems often spell identity differently — one holds first_name and last_name, the other full_name. A key is therefore a small recipe rather than a bare column list: a part can be a column as it is, concat (join columns with a separator), split (take one piece of a delimited column), substring, digits (drop punctuation from an account number) or pad (left‑pad to a fixed width). concat(first_name, last_name, ".") is one part and pairs with column(full_name). Splitting is offered but marked lossy on the row — nothing can rebuild "van der Berg" from a space‑delimited name — so joining both sides to one value is the shape to prefer.
Normalisation is stated, never silent. After the recipe the join is still exact: JOHN.SMITH is not john.smith. Each key carries visible toggles — ignore case, collapse spaces, drop punctuation, fold accents — defaulting to case and spaces, and the run log prints the recipe and its normalisation beside every table. Choosing none says exact.
6.9.3 Preview keys, and Suggest a key
Preview keys reads a few hundred rows from each table, builds the keys, and shows real values side by side with the overlap between tables, keys that repeat within a table, and a verdict — before the set is saved. A low overlap is a warning, not a refusal (the systems may genuinely hold different customers, and those rows report as coverage gaps). A repeated key blocks the save only on a table that has asked to refuse repeats (§6.9.4).
Suggest a key measures every column pairing across the first two tables on the same sample — how unique each candidate is, how much it overlaps, how often it is blank — and ranks them, overlap first: cust_id against customer_no can be perfectly unique on both sides and share not one value, and a pairing that never meets is not proposed. When AI assistance is configured, the model reads the ranking and the column names and says which one a person should pick and why, warning when the best‑scoring key is personal data used as an identifier. Only column names and the measurements leave the server; the ranking stands on its own without a key; and the advisor proposes — Apply is a button a human presses.
6.9.4 Duplicates: distinct, and repeats within a source
Two rows in one table can share a key for two different reasons, and only one is a problem with the key. Distinct (a tick per table) collapses rows that repeat the key and hold identical values in every compared field — one record written twice. It never collapses rows that disagree: choosing between them would be a survivorship decision a read‑only comparison has no right to make.
By default a repeated key is reported, not refused: the run proceeds, every copy takes part in the diff, and copies that disagree are counted as a conflict within that source — the result names the source, and the grid shows every value it holds, tagged as a conflict inside it. A run that found such repeats says so above its tiles ("12 record keys occur more than once inside a single source … 8 of the 12 conflicts below are of that kind"), because a CRM holding one customer twice under two ids is a finding, and usually the first thing to fix. A table whose key is meant to be a strict identifier can be set to refuse repeats, in which case the run stops with the repeated value and its count.
6.9.5 Reading a run
A run reports records compared, field agreement, and the four status counts, and it says three more things when they apply: that it was capped (a partial extract, with the cap named); that distinct collapsed rows (how many); and — most importantly — that nothing was compared. Zero conflicts is good news only when rows were paired; when no key on one table ever met a key on another, the run says so, tells the two apart (no rows · rows but no usable key · one source only · keys do not join), and lists how many keys each source produced with an example of each. The figures underneath are then zero because the comparison never happened, and the screen says exactly that. A set that ran but paired nothing reads ran · nothing compared in the list, never not run.
7. Governance
Governance controls how change happens and keeps a defensible record of it.
7.1 Policies
Role: Administrator to create and edit; every governed role can read.
The Policies tab holds your governance policies and shows the overall governance posture. A policy is a written commitment with a target attached — "customer data quality shall be at or above 95%" — against a domain or the tenant.
Only two policy types are measured, and you should choose accordingly. The platform computes a live current value for exactly two kinds of policy:
- Data quality — measured from the quality score for the scope.
- Exception resolution rate — measured from exception throughput.
The other policy types the form offers are recorded, published and audited, but nothing computes a current value for them. They are statements of intent with a number beside them, and no refresh will ever fill that number in. If you want a policy the platform will hold you to, write it as one of the two above.
The Policy Advisor proposes policies from your own data, and it is deliberately narrow for the same reason: it proposes only the two measurable types, because it refuses to propose a target that nothing can check. It also skips any domain whose quality has never been computed — a domain with no measurement is not a domain scoring zero — and it names every domain it skipped and why. Proposals arrive pending; accepting one creates the policy exactly as the Policies page would, through the same approval gate. Each run is capped, and the advisor is deterministic: the same evidence produces the same proposals.
7.2 Change Requests
The Change Requests tab is the review‑and‑approve workflow. Governed changes are submitted as change requests, reviewed by an authorised approver, and only applied once approved. Every governed write in the platform passes through one gate, so there is no side door.
What can be governed. Not every action is reviewable — the governable set is fixed, and covers matching rules, survivorship, DQ rules, standardization, managed values, governance policies, creating a glossary term, updating a catalog facet (§3.8), deleting a domain, entity or relationship, and creating or updating a campaign. Which of those actually require review is configurable per area in Access Control (§10.3).
The list is grouped by request type and follows the collapse rule in §2.10 — the grouping comes from a fixed vocabulary, so a quiet request type is folded rather than absent.
The request form states what approval will actually do. Per artifact type, the form says whether approval creates the change automatically (green) or hands it back for a person to finish in its module (amber). A hand‑back is a designed outcome, not a failure — the approval is the permission; the module is where the work happens.
Targeted items — what this request actually touches. A request carries a list of the objects it changes: the domain, entity, attribute, DQ or standardization rule, match ruleset, survivorship rule, reference code set, mapping set, ingestion job or report the change lands on. That list is what the printed change record shows, what an artifact's own page reads when it warns you that an open request already touches this object, and what an implementer works through — each item is ticked implemented on its own, so a request touching six things can show five done and one outstanding rather than a single all‑or‑nothing state.
You browse for the object; you do not type its name. Adding an item opens an objects navigator — the model itself as a walkable tree, rooted under the active model's name with a MODEL badge: domains, the entities beneath them, and beneath each entity its attributes, relationships, DQ rules, match rulesets and survivorship rules, with Standardization, Reference, Integration and Analytics as sibling branches. Typing in the search box opens matching branches as a view — your own expanded/collapsed state is restored exactly the moment you clear the search, the fold rule from §2.10. Because you pick a real object, the link carries that object's id and its live name: rename the entity next month and the request still points at it, and still reads correctly. A hand‑typed label was checked against nothing and pointed at nothing.
Something that does not exist yet is still allowed — and is marked as such. The entity this request is asking for can be added by name from the same panel. It lands unwired: no id, and the request says so rather than implying a link it does not have. Once the artifact has been built, wire the placeholder to the real object and the promise becomes a link.
The panel is honest about what it knows. What has already loaded stays on screen and stays clickable while a search runs, so the list never blinks out from under your cursor. A hub that does not answer in time says so, rather than rendering as an empty tree — an empty result reads "nothing to browse" only after a load that actually succeeded. And a hub too old to have the browser at all says exactly that, naming the backend build it needs, instead of failing as a mystery.
Every change to the list is in the request's event trail — including removals. Linking an item, wiring a placeholder to a real object, and removing an item are all recorded against the request, each naming the item concerned. Removing one narrows what an approver is being asked to approve, which makes it the least acceptable of the three to record silently.
Two related guards worth knowing: deleting a mapping set with dependent ingestion jobs refuses and names the jobs — with staged‑record and run‑log counts — and requires an explicit cascade confirmation; and write‑back Update operations can be put behind workflow approval via the Access Control matrix (§10.3) — that switch ships off, so nothing changes until an administrator deliberately turns it on.
Signing the change record. The printed record ends with four parts — Requested by, Reviewed by, Approved by, Implemented by — each already carrying a name and a timestamp written by the governance lifecycle from the signed‑in session, with a second copy in the audit trail. On top of that, the person who performed a part may confirm it: an attestation that they read the record and stand behind it. You may confirm only the part you performed. Requested is not signable at all — raising the request is the assertion, and asking for a second confirmation of your own request adds nothing — and a part the platform performed reads "applied automatically" rather than leaving a line a person could claim.
A confirmation records which version was read. The record is generated on demand from live rows, so a request returned for rework and re‑approved renders differently under the same four names and times. Each confirmation therefore stores a fingerprint of the governed facts as they stood at the moment it was given — title, artifact and change type, description, justification, domain and entity, the targeted items, and the four lifecycle names and times. Re‑word the justification afterwards and the panel shows that confirmation as signed against an earlier version, asking the same person to re‑confirm. Re‑print the same record on a different day and nothing moves: the fingerprint is taken over the facts, never over the rendered page.
Three states, never two. A part reads confirmed, signed against an earlier version, or awaiting confirmation — silence and consent must not look the same, so an unconfirmed part says so rather than printing a blank line. Seeing that a part is confirmed is not a privilege: everyone who can open the request sees the panel, and only the eligible person sees an enabled button — one that, where it is disabled, carries the sentence saying why. Each confirmation also carries a short verification code, shown beside it and printed on the record, so a paper copy can be traced back to the confirmation it came from.
7.3 Review bypass (suspending governance review)
Review bypass is a per‑model switch that temporarily suspends governance review. It has two states, and the studio is deliberately loud about one of them:
- Review bypass OFF — normal governance protocols. Governed changes go through the change‑request workflow (§7.2). This is the default and the state a live model should be in. Nothing appears in the header — a plain header is the statement that governance is running. In Deployment Models the model's row shows a quiet shield icon; that icon is how you turn the bypass on.
- Review bypass ON — governance review is suspended. Governed changes apply directly, the moment they are made, with no approval step. This is the natural mode during initial model build. The header turns amber for everyone working in that model, an amber Review bypass ON pill sits next to the checkout control (with a Turn off action), and the model's row in Deployment Models carries an amber Bypass ON chip.
Where each direction lives. Turning the bypass on is deliberate and lives in one place — Administration → Deployment Models, on the model's row. Turning it off is available from anywhere, on the amber header pill, because the state you want to leave should never be hard to leave.
The amber marks the suspended state, not the working one. A plain header means governance is doing its job; amber means it has been switched off and somebody needs to switch it back.
Turning the bypass on or off always asks for confirmation, walks the model checkout gate where required, and is audited (GOVERNANCE_BYPASS_ON / GOVERNANCE_BYPASS_OFF). Turn it back off once initialisation is done, so normal approval resumes.
Be clear about what "bypass on" means. It is not a lighter review — it is no review. It is a per‑model switch, held by administrators and developers, and every flip of it is in the audit trail. A production model left with the bypass on is a decision somebody made, and the audit trail will say who. Changes applied while the bypass was on stay applied when it is turned off — they are not retrospectively sent for review.
7.4 Model checkout
Model checkout (see §2.4) complements governance by ensuring only one editor changes a model's structure at a time.
7.5 Strict pipeline mode
Role: Admin or Developer; per model; audited.
Strict pipeline mode makes the pipeline refuse to cut corners:
- Survive requires an active match ruleset with a recorded run.
- Publish requires active DQ rules with no open critical exceptions.
When a run is blocked, it fails loudly with an explanation panel and deep links to exactly what's missing. An Override & run escape hatch exists for genuine emergencies — every use is audited. Toggle strict mode per model in Deployment Models.
7.6 Approval Workflows
Role: Admin, Steward, Data Owner or Developer.
A change request needs one approver. Some changes need three, in order, from different roles. Governance → Approval Workflows is where you build those:
- A workflow is a sequence of steps. Each step names the role that approves it and the SLA it must be answered within.
- Approve advances to the next step; the last approval applies the change. Reject routes the request back — where it goes is part of the workflow's definition, not a fixed rule.
- Every action is kept as history on the request: who, when, which step, what they said.
Workflows govern two things today: change requests (§7.2) and authored golden records (§5.5). A step waiting on your role appears in My Work as a workflow approval, ranked by its SLA.
7.7 Who is consulted and informed (RACI)
Approval answers who said yes. RACI answers who should have been told.
MDM Studio reads RACI from the catalog (§3.8). Each asset carries four facets:
| Letter | Facet | Effect on a governed change |
|---|---|---|
| A — Accountable | Owner | Named on the request. Inherited from a governing glossary term if the asset does not name one. |
| R — Responsible | Steward | Named on the request. Also inherited from a governing term. |
| C — Consulted | Consulted | Listed on the request, so a reviewer can see whose opinion this change was supposed to attract. |
| I — Informed | Informed | Notified when the request is raised. |
Only A and R are inherited from a governing term. Consulted and Informed are set on the asset itself, deliberately — being told about a change is a specific arrangement, not something to inherit by accident.
The snapshot is an audit fact, not a live view. At the moment a governed change is raised, the platform writes a raci_snapshot onto the request and a matching history entry. It records who was accountable, responsible, consulted and informed then. Change the facets tomorrow and yesterday's request still shows yesterday's people — which is the whole point when someone asks, six months later, who was supposed to have been consulted.
Names that could not be notified are reported by name. If an Informed facet holds a person who is not an active user, the request card says so explicitly and lists them. A silent failure to notify would be worse than no notification at all, so the platform refuses to be silent about it.
The As‑Is Assessment (§3.6) has its own, separate RACI: a letter per stakeholder and a matrix per planned activity. It describes the programme; this one describes changes to data.
8. Analytics
The Analytics hub turns mastered data into insight.
8.1 Dashboards
The Dashboards tab lists saved dashboards; the Dashboard Builder lets you compose them (it opens from a dashboard, not from the tab strip). The Home page is itself an operational dashboard (golden‑record volume, data quality, open exceptions by severity, ingestion jobs, domain health and recent activity).
The builder is a canvas, not a fixed grid. Drag a report from the toolbox or drop it on the canvas; each tile carries its own fit behaviour (reflow, fit, fill, actual size, anchor, lock aspect), and the canvas itself takes a colour, an image, a pattern, a grid and a zoom. Tile titles are formatted as text — size, weight, alignment, font — rather than configured as settings. The toolbox groups reports by area and remembers which areas you collapsed; typing in its filter opens every area that matches, so a search result is never hidden behind a fold you set last week.
A tile is moved and resized from anywhere on it. A press anywhere on a tile begins a move — a short drag threshold keeps a plain click as a selection — and every tile carries eight resize handles rather than a single corner, so whichever edge is on screen will do. Arrow keys nudge the selection and shift‑arrows resize it, which is the only thing that still works when a tile has scrolled entirely out of view. None of this depends on the tile's header strip: a tile whose header the author hid, whose top has scrolled above the canvas, or which sits beneath another tile is still fully editable.
8.1.1 Filters that span the whole dashboard
Tiles on one dashboard usually come from different sources, and the same business idea is rarely the same column name in all of them — a region might be region_name in one source and sales_region in another. A dashboard filter is therefore a concept with a binding per source: you name the idea once, and say which field each source uses. One control then drives every tile that has a binding.
Three things this deliberately does not do:
- It does not guess from names alone. The unification detector proposes a binding only where the candidate fields genuinely correspond, and it states its confidence. Where a name is shared but the vocabularies are not —
statusacross five sources that each mean something different by it — it refuses to unify rather than silently filtering the wrong rows out of a tile. - It does not quietly leave a tile unfiltered. A tile whose source has no binding for the filter is handled by a policy you choose per filter: flag it (the default — the tile says it is not covered), ignore it (it keeps showing everything, knowingly), or blank it. Flagging is the default because a tile showing unfiltered numbers next to filtered ones is the failure that is hardest to spot.
- It does not replace the report's own filters. A report keeps whatever filters it was authored with; the dashboard filter is applied on top, and the tile shows both.
Where the names do not match, you propose the filter yourself. The detector groups by exact field key. That is the right rule for something offered automatically, and it is also why the panel so often reports nothing left to unify: customer_id here is cust_id there and CustomerID on the third, and the one filter the author actually wants can never be suggested. Propose one by hand opens a panel where you name the idea and choose the field on each source; it offers a likely field per source and states why it thinks so, and it will not create the filter until at least two sources are bound — one binding is a per‑tile filter wearing a different hat. A proposal is labelled yours, not detected, and it is never applied automatically at any similarity score: two columns with similar names can hold unrelated things, and a unified filter that binds them sends one value to both, so the tile on the other vocabulary comes back empty and reads as "no data" rather than "wrong filter".
8.2 Report Studio
Report Studio is a full authoring environment inside the hub. It builds reports over vetted, governed data sources — golden records, data‑quality scores and exceptions, match results, reference data, pipeline runs, governance requests and registered custom queries — and it reads them through the product's own access layer rather than around it.
8.2.1 The visual set
A report is one visual. There are seven, and switching between them brings the wells that visual needs:
- Table — many records, with drill‑through and per‑column formatting.
- Record — one record with everything related to it.
- Form — a single record, laid out as fields.
- Pivot matrix — rows × columns × values.
- Pivot — grouped rows with subtotals and a drill ladder.
- Chart — column, bar, line, area, pie, donut, scatter, radar, treemap, funnel and waterfall.
- Scorecard — KPI cards, optionally split by a dimension so the card set repeats once per value.
A chart can mix marks on one axis — plot one measure as bars and another as a line — and any chart or scorecard can be split into small multiples, one panel per value of a dimension, drawn on a shared scale so the panels are comparable.
8.2.2 Authoring: fields, canvas, properties
The studio is three panes. Fields lists what the source offers; the canvas holds the wells — Rows, Columns, Values, Filters and Small multiples (spelled Split by on a scorecard) — and Properties configures whatever is selected. Drag a field into a well and the live preview redraws. The whole session is undoable and redoable, and ⌘S / Ctrl‑S saves.
Measures aggregate with Sum, Average, Count, Distinct count, Min or Max, and carry ordering, a Top N with an optional Other bucket, having‑style thresholds, and a Show as mode: the value, % of total, % of a grouping — the share within any field the report groups by, including the Split by field of a scorecard — running total or rank. The list of groupings is built per report: a share of a field the report does not group by cannot be computed and is not offered, and if a drill later removes that field the mode says it no longer applies and falls back to % of total rather than silently answering a different question. On a matrix the classic % of row / column / grand total remain.
Two behaviours worth knowing because they protect the number on screen: a subtotal over an average is weighted by the records behind each group, not averaged again (an average of averages is not an average); and a distinct count refuses to be re‑aggregated at all, because distinct counts cannot be summed.
A date dimension can be grouped by period. A date field in Rows or Categories carries a Date grain — raw value, day, week, month, quarter or year — so a trend is not one point per instant. The truncation is performed by the hub's own engine, and the server echoes back the grain it actually applied, so an axis cannot be labelled with a period the query did not group by. A source that lives on another database is left ungrained rather than guessed at, and the control appears only on date fields.
The properties pane folds. Every group on both the Data and Format tabs is a collapsible section, with Expand all and Collapse all above them, and they open collapsed — the compact view an author who knows the panel actually wants. Two things stop that from hiding the controls: a folded section states its contents in its own header (a Totals section reads outer subtotal · grand total · row totals when shut, or none shown), and a well unfolds the moment a drag enters it, so a collapsed drop target never becomes an unreachable one. Your own folds are remembered per section, and a choice you made always beats the shipped default.
Recent reports. The report picker opens with a Recent list above the groups: the few reports you actually work in, each with when you last opened it and how many times. The order is frequency with recency decay rather than either alone — the report you open every morning does not drop off because you spent an afternoon elsewhere, and one you opened five minutes ago still outranks one you used heavily last quarter. It is kept per user, and it holds report ids only: every row is resolved against the list the server returned for you, so a report that was deleted, renamed away, or is no longer yours to open simply is not there. Rows can be dropped from the list individually, the list stands aside while you are searching, and the groups below still carry the same reports — it is a shortcut on top of the tree, never a replacement for it.
8.2.3 Filters
One operator vocabulary applies across every visual. A filter can be marked required, exposed as a parameter, or given a default, and the reader is shown which. Relative dates — last 30 days, this quarter — resolve when the reader opens the report, not when it was authored.
Filters become pickers automatically. When a field has few enough distinct values, its filter renders as a dropdown instead of a text box. "Few enough" is a single central setting (Filter dropdown threshold, §2.6.1) rather than a per‑report choice, so the whole product behaves consistently. The distinct values are read through the same governed door as the data itself, so masking, row‑level access and tenant isolation all still apply — a picker never reveals a value you could not otherwise see. Where a field has too many values to list, the control says so and names the count.
Typing is a search, and it says so. On a text filter that is not a picker, what you type is applied as contains, not as an exact match — the control's placeholder states this. Authors who need exact matching set Match exactly on the filter.
8.2.4 Calculated fields
A calculated field is an expression saved on the report and used like any other field. Expressions are compiled to SQL for whichever hub engine you are on — nothing an author typed is passed to the database as text — and the compiler is the same one the data‑quality rule authoring surface uses, so an expression that is legal in one place is legal in the other.
8.2.5 Formatting, and the marks that refuse
Values honour the display format governed on the attribute in the Data Dictionary (§3.4) — in grids, in chart data labels, on axis ticks and on KPI cards. Formatting is display only: filters, copy, CSV, generated SQL and every API payload carry the raw value.
Column tools: one rule set, several marks. A column can carry a tool — a mark drawn in place of the plain value. There are five: an indicator (a dot or icon stating which rule matched), a heat bar (a proportional bar), a badge (the value in a tinted pill), a range (the value inside an expected band, against a target) and a delta (movement against another column). They are not five features. They are five ways of drawing the same rules, and the rules belong to the column: change a threshold and every mark on that column moves together, so a bar and a badge on one column cannot disagree, because there is only one place the answer lives. The same tools reach matrix and pivot measures, and a reduced pair — indicator and badge — reaches row and column headers, matched against the header's own text. A column with no tool keeps every existing behaviour untouched.
Rules are typed, and they are not only for numbers. A rule is a condition, a tone, an icon and a label. Numeric operators (≥, >, ≤, <, =, ≠, between, outside) sit beside text operators (is, is not, contains, does not contain, starts with, ends with, is one of a comma‑separated list, and a wildcard matches using * and ?) and date operators (on, before, after, between, and windows counted in whole days from today — in the last …, in the next …, older than … — so a rule written today is still true next quarter). Text is matched without regard to case. The editor floats the column's own family to the top of the operator list but does not hide the rest: a text operator on a numeric column ("code starts with 4") is legal and sometimes exactly right. A rule's value can also name another column, written @column_name, which is what turns a fixed threshold into a comparison the row carries with it — each domain judged against its own target rather than one number for all of them. Rules are read top to bottom, the first that fires wins, and otherwise is the last resort.
Blank is its own answer, not a failed comparison. A blank cell falls through every operator except is blank, and is never coerced to zero. An unmeasured quality score painted red as "0 % — at risk" is a different and much worse claim than "not measured"; the blank / not‑blank pair is how an author says which one they mean.
Colour never carries meaning alone. Every rule carries an icon and a label as well as a tone, so a reader who cannot separate your green from your amber can still read the mark — the default tones were checked against the application surface for contrast and for colour‑vision deficiency, which is a band that requires a second encoding. Tones are chosen from the palette the rest of the product uses, not from an operating‑system colour dialog.
The heat and range axis is measured, not assumed. By default a bar is fitted to the column's own range, with the floor anchored at zero whenever the data is non‑negative — a bar that starts at the smallest observed value exaggerates every difference above it. An author who wants a stated scale rather than the data's — a 0–100 quality axis that stays 0–100 in a month when nothing scored below 90 — sets a fixed minimum and maximum instead. A diverging blend is offered, and it is right for "above or below target"; it is deliberately not the default, because a midpoint that is merely the middle of the axis renders a whole column grey.
Totals are decided separately, and marks reach them only if you say so. A matrix chooses its subtotals (every level, outer level only, none) and whether the grand‑total row and the row‑totals column appear at all; then, as two further opt‑ins, whether the measure's marks apply to total cells, and whether the header rules apply to the "Total", "Grand total" and subtotal labels. A pivot carries the same pair for its total row. Marking totals is off unless asked, and the panel says why: a rule reading "≥ 95 healthy" is a claim about a percentage, while the total beneath it may be a sum whose scale has nothing to do with 95.
Depth. A report's Look is flat, 3D raised (lit from above) or 3D sunken (pressed into the page), and the marks follow it: a pill, a bar or a badge takes its depth from the report unless you set the mark's own — follow the report, flat, raised or sunken. A heat bar's track stays inset under both 3D looks, because a bar reads as something filling a channel; what raised and sunken change is the fill. Depth is drawn on the mark itself rather than by stylesheet, so it survives the PDF and PNG render path, and a report left flat renders exactly as it did before these controls existed.
Table columns also take the older conditional formatting — colour scales, icons and value rules — which is untouched by the tools above and still renders on every report that uses it. KPI cards render as plain values, bullet bars or gauges, with bands; where a measure is scored lower‑is‑better, the bands paint in that direction.
The band editor. Bands are set on one control: a strip showing the three regions (bad · warning · good) with two draggable handles, and four inputs — Bad below, Good from, Scale min, Scale max — that agree with it. The inputs are unit‑aware: under % of total they read in percent while the measure stores a fraction, so "good from 70" means 70 % and not 7,000 %; under rank they read as positions. Changing Show‑as converts the bands and says so — fraction ⇄ raw units are rescaled with a note; rank clears them, because a position is not a quantity. Bad‑below cannot exceed good‑from. Beneath the strip is the real range of the data the card aggregated (min · median · max), and the editor warns when a band can never be reached — a "good from" above every value on the report marks every card bad, and a threshold no value can cross is not a stricter standard but a broken one.
Suggest from data reads the shape of the distribution rather than filling the quartiles blindly: on an even spread it proposes the quartiles; on a long tail — which is what a percentage of total almost always is — it proposes the median as the good line, so good means better than most rather than better than nearly all; on flat data it says there is nothing for a threshold to separate. It fills the boxes and applies nothing. When AI assistance is configured, one sentence of reasoning in the measure's own vocabulary appears beneath the arithmetic, labelled as the model's; without it, the arithmetic is the whole answer and the screen says so.
Frame, background and panel borders are the author's too. Under Format ▸ Frame & background the report's border takes a colour, a line style (solid, dashed, dotted, double), a thickness and a corner radius — or is hidden — and the report background is a colour or none, so a report can sit flat on a dashboard with the page showing through. A chart's small multiples and a split scorecard's panels have the same controls per panel: border, background, and titles that follow a template ({field}: {value} is the default), a text of your own for a particular value, or no title at all. Small multiples can also open with a Total panel — the un-split figures by category, read as its own query and never summed from the panels, because an average of per-panel averages is not the average. Anything left unset renders exactly as it did before these controls existed.
Reference values are the reader's choice, and it is remembered. A field bound to a code set, a value hierarchy or a tier can read as its display value (South Africa) or its stored code (ZA). Each report carries its author's choice under Format ▸ Reference values; every reader can override it for themselves — the small Report · Display · Codes pill on any report, or My Account ▸ Reference values in reports — and the choice is saved per person and applied to every report, dashboard, filter, drill and export from then on. Golden keys, wherever a report shows them, are trimmed to the system segment and the first digits of the digest, with the full key in the tooltip.
A gauge or bullet resolves its scale through a stated ladder — authored bounds, then a shared scale, then the value's format, then the target, then the previous value, then the data — and prints the ends it chose, so a reader can see what the mark is measured against rather than inferring it.
Some data cannot be drawn honestly, and in those cases the visual refuses and says why instead of drawing something wrong:
- A pie, donut or treemap cannot represent a negative as a share of a whole.
- A funnel whose stages are all zero has no shape to show.
- A scatter needs two measures; with one there is nothing to correlate.
- A value that could not be computed renders as unavailable with its reason, never as a confident zero.
8.2.6 Saved views, export and subscriptions
Any set of filters, sorts and drill positions can be saved as a named view, made your default, or shared with the report; a view is addressable as ?view= on the report's URL. CSV and Excel exports are generated on the server from the same governed query the screen ran, so an export is the report rather than a re‑derivation of it, and print keeps the layout.
A report or a whole dashboard can be subscribed to on a schedule and delivered by email. Each run executes as the subscriber, so a subscription can never deliver rows the recipient could not open themselves. Where no mail server is configured the run is recorded as SKIPPED rather than reported as sent.
Subscribe ▸ opens the manager for that report or dashboard: each subscription is a card — its name (or recipients), the schedule in your local time with the UTC it fires at, formats, the send condition and the saved view, whether it is active or paused, the last run's outcome (SENT, SKIPPED or FAILED, with the reason) and the next run. Run now answers on the card itself. A subscription can be edited in place — one or many recipients, a name, a note that rides in the mail, every weekday / day / week / month or a custom set of weekdays, the time (entered locally, stored in UTC; the form says when the two fall on different calendar days), formats, condition and view — paused and resumed, its run history opened, or deleted. Every subscription you hold, across all reports and dashboards, is also listed under My Account ▸ My subscriptions. A subscription can also be sent by event instead of by clock: when a named job, job group or execution package completes, fails or starts, the report goes out moments later with the data as it stands then. PDF and PNG are rendered on the hub: a headless browser opens the report as the subscriber — same access scope, same masking — and prints it; where that browser has not been installed the two formats cannot be selected, and the manager names what is missing.
Separately from report subscriptions, jobs, job groups and execution packages can notify you of starts, successes and failures, in‑app and by email (§4.4.1). One incident produces one message: where you are subscribed at more than one level, the most specific scope wins, so watching both a job and its group does not send the news twice.
8.2.7 What a report may read
Analyze in Report Studio on a golden‑records view starts a report preloaded with that entity's golden data source.
Sensitive fields above your clearance are masked in report grids and cannot be used as filters, group‑bys or measures (this prevents inferring hidden values). Exports inherit the same masking. Filter and sort keys are allow‑listed against the source's own catalog — a key that is not in the catalog is refused by name rather than passed through — and authoring custom SQL sources requires Developer or Admin.
One limit to be aware of when you design a report. Masking is applied to reports built over entity‑stage sources — golden records and the model's own data. A report built over a raw pipeline stage source is not masked by the report layer. Treat raw‑stage report sources as an engineering tool and publish entity‑stage sources for general consumption.
Some built‑in sources are access‑group scoped in their own SQL. Catalog sources that reach golden records through a domain or through the exception book — the domain overview and exceptions‑by‑severity among them — carry the caller's access‑group predicate inside the query rather than filtering the answer afterwards, so the counts they return are counts of the population you may see. Those sources fail closed: resolved without a user context, or with a scope that cannot be resolved, the request is refused and says so rather than being answered unscoped. Where your access is genuinely unrestricted the predicate contributes nothing and the source behaves exactly as it always did.
8.3 Report Viewer & Pipeline Health
- Report Viewer renders a saved report for consumption. It opens from a report rather than from the tab strip.
- Pipeline Health shows the health of your ingestion and lineage pipelines. Alongside stage health it carries a source‑to‑target mapping view, which unifies your field mappings and attribute sources into one lineage statement per target attribute, and a baseline‑versus‑post‑MDM comparison that shows what mastering actually did to quality. Source‑to‑target mapping is empty until attribute sources and field mappings are registered.
External specifications are counted separately, and never blended. Where mapping sets target external assets (§4.2), the source‑to‑target view carries them as their own External specifications section, with its own list and its own counts. They are deliberately kept out of the attribute‑coverage numbers beside them: an external specification does not cover one of your attributes, and adding it to a coverage percentage would make the number mean two things at once. Each row deep‑links into the Mapping Designer with that set already open.
One reading rule for that section: on a hub whose specification version tables predate this release, a set's version state shows as NOT MEASURED, not as zero. Those are different facts — nothing to report yet is not nothing there — and the section says which one it is.
8.4 GR Explorer
Role: any role that can view golden records. Read‑only.
Analytics → GR Explorer walks the golden set through its relationships rather than one entity at a time. It is built for the question "show me this customer, and everything hanging off them".
- A relationship map dropdown lists every domain‑to‑domain link in the model, with the number of relationships in each. Pick one to focus that domain; the relationships it covers are then printed underneath — from entity → to entity · cardinality — rather than hidden in a tooltip.
- The main grid is the governed golden‑record grid for the focused entity, with the publish line, the population badge and exception‑ordered columns.
- One tab per relationship renders the related records, resolved through the relationship's own key columns — orders for this customer, addresses for this party.
- A count on every relationship tab says how many related records this record actually has, so you can see which relationships hold anything without opening each one in turn. A count that could not be computed draws a dash, never a zero — none related and this relationship cannot be resolved are different answers, and only one of them is about the data. A grey 0 means the server really did say zero, which is usually the answer you were scanning for.
- Pivot: click a related record and it becomes the new focus, leaving a breadcrumb trail behind you — Customer ▸ Order ▸ Product. Click back along the trail to return.
The shared golden row menu (§2.10) is available throughout, so you can trace, inspect or copy a key from wherever you have arrived. The related grid has its own right‑click menu, fully scoped to the related entity — exceptions, trace, lineage, sample values, filter (with a visible, clearable chip) and copy key.
The key search can reach through the relationships. With a term in the search box a related keys toggle appears beside it. Turn it on and a golden key belonging to a related record matches too, so a Sales Order key typed while you are on the Customer entity finds the customer that owns it — which previously meant switching entity, reading the join column off the row, switching back and filtering by that value. The toggle is offered only while there is something to search, because a control with nothing to act on is a control with no effect.
Two properties are what make the result safe to quote. First, it says how far it reached: the grid carries a chip naming the related entities that were searched, and names the ones that were skipped with the reason — the relationship declares no attribute on both sides, or you have no access to that entity — and where a model carries more relationships than one query should correlate, the cut is stated rather than hidden. A silently narrowed search is indistinguishable from a record that is not there. Second, a record you are not cleared to see can never be used as a lens on one you are: matching through it would itself disclose that a record with that key exists and is linked to this one, so the search fails closed on the related record's sensitivity instead of filtering the row out afterwards. Searching by a restricted record's key therefore finds nothing through it — the correct answer, not a gap to be closed later.
The Hierarchy rail. Entities with a navigable hierarchy gain a Hierarchy control that opens a rail beside the grid, with a source picker offering up to four never‑blended sources:
- Tiers — the model's tier dimension as a ladder. It is bidirectional: clear buttons per pick, up/down steppers, arrow keys; clearing a broad tier drops the dependent narrower picks (the tooltip warns first), and a stepper refuses at the ends with a stated reason. Where a tier's field is standardized but not yet mastered, the ladder notes that it groups by the cleansed value (§3.7).
- Record tree — the entity's own parent‑key structure in the published data, shown with business names (name leading, code muted). Unparented records group under a visible amber "Unparented (N records)" divider instead of masquerading as roots.
- Managed — a maintained hierarchy from the Hierarchies workbench (§3.7).
- Value hierarchy — a declared containment chain bound on the entity (§3.10), walked level by level: pick Gauteng, then Johannesburg. The deepest pick is the filter — the narrower value is already inside the broader one, which is the whole difference from tiers, where the levels are independent columns and every pick is sent. An optional derive toggle (off until you turn it on) also shows records placed by the hierarchy from a lower level; they are labelled, and nothing is written.
Picking a rail value filters the grid, including cross‑entity (a Product Category tree counts and filters Products); "No value at this tier" and "Not in this tree" buckets are always shown with counts. The rail is a persisted treeview — one fetch for the whole tree, expansion remembered per source and entity, an oversized tree says it was capped, and a capped branch is never drawn as a leaf. The node label attribute is editable in the rail itself (§3.7). Where two sources disagree about the same records, the disagreement is surfaced as a finding, never merged away.
Rows that can be walked carry a ⤺ walk affordance in the leading actions cell (always visible, whatever your column order); it opens the tree navigator focused on that record — ancestors, siblings and children as lanes, any of which can become the new focus. A hierarchy‑filtered grid that comes back empty names the filter that emptied it, in the model's own words, and offers to clear it.
Every empty grid says which of the possible reasons applies. A search that matched nothing says so and offers the count you would go back to — naming the related entities as well, when the wider search ran. A hierarchy filter names the branch. And an entity whose counts show published records is never told to publish its pipeline, whatever combination of search, filter and view produced the empty result: "No golden records — publish this entity's pipeline first" is reserved for the case where it is actually true. A searched record that simply does not exist is a different fact from an entity that was never published, and the screen distinguishes them.
An entity with no hierarchy shows no Hierarchy control at all. Absence is the honest answer — there is no greyed toggle to wonder about. If the entity's data merely looks like a tree, the Explorer may show a measured declare‑this‑hierarchy suggestion instead (§3.7).
Layout is yours: Stacked or Side‑by‑side, a draggable divider, either half hideable — all remembered per user, independently of the rail.
The rail is not only here. The same rail, with the same sources and the same laws, mounts on the Exceptions workbench (§6.3), the pre‑publish Golden Records browser (§5.2) and the Golden Fields records grid — one implementation, so a branch means the same records on every surface, a capped scope blocks everywhere, and every count shown beside a scoped grid counts the scoped population.
GR Explorer reads published golden records. Until an entity has been through Match → Survive → Publish it has nothing to show.
8.5 Report Center
Analytics → Report Center is the gallery: published reports and dashboards, presented for people who consume them rather than build them. It is where you send a business audience — no data sources, no expression language, just the reports somebody chose to publish.
The gallery is empty until report definitions exist. That is the normal state of a new deployment, not a fault.
9. Reference Data (RDM)
Reference data is the governed, shared vocabulary of your organisation — code sets like country codes, currencies or status values.
- Governed code sets with immutable releases: a release is a frozen, versioned snapshot that downstream systems can rely on.
- Crosswalks map between different source vocabularies and your canonical codes.
- The RDM Serving API (Administration → RDM Serving, §10.8) exposes released reference data through a read‑only API secured by revocable service tokens, with checksum‑based polling so consumers only fetch when something changed. Stewards and data owners can see the serving page read‑only; token administration stays with administrators.
- RDM Integration (Administration → RDM Integration, §10.17) is the push half: downstream systems subscribe, and a signed webhook tells them a new release exists rather than waiting for them to poll.
Day‑to‑day maintenance of reference domains and managed values happens on the Integration Hub's Standardization page, on its Reference Data and Managed Values tabs (§4.3) — the former standalone pages redirect there. Every reference domain and managed value shows where it is used (mappings, standardization rules, jobs), and destructive changes are guarded by a usage breakdown.
Display values are the standardization. A code value carries a display value and any number of aliases, and the platform folds every spelling it meets — sa, za, ZAF, South Africa — onto that one display label through a single shared fold. The hierarchy rail, the tier ladder and Data Quality's By‑hierarchy lens all group by that label, so one country never appears as four nodes, and a chart's legend and labels are placed to avoid colliding with each other. An alias whose canonical code is not live resolves to nothing rather than to the raw value: handing back a code nobody approved is worse than leaving the value alone.
Reference data is a shared, org‑wide standard: every tenant's models map to the same governed reference library.
10. Administration
The Administration hub is the platform's control centre. Its tabs are grouped as Identity, Security, Monitoring, Storage, Reference data, Platform and Account. Most require the Administrator role; some also require platform superuser; a few (Users, Roles & Permissions, Performance, Scale Readiness, and the reference‑data tabs) are open to data owners or stewards as well, and My Account is open to everybody.
If you are a steward, data owner or analyst, this hub is inside the collapsed Engineering disclosure in the sidebar (§1.3).
10.1 Tenants
Role: Administrator and platform superuser.
The Tenants tab manages the isolated workspaces of the platform:
- Create and edit tenants; set per‑tenant branding (logo, badge colours and font). Reset style clears and saves immediately.
- List and manage users under a tenant, including adding an existing user from any tenant (a user can belong to more than one tenant).
- Suspend a tenant — full suspension blocks all sign‑in and access for that tenant until reactivated.
- Manage platform superusers. Superusers come from two sources: a configuration list (bootstrap/break‑glass, set by the deployment) and granted superusers (toggled in the app by existing superusers). The platform prevents removing the last effective superuser.
10.2 Users
Role: Administrator or Data Owner.
The Users tab manages accounts within your tenant:
- Create users with a username, email, display name and role.
- Choose the sign‑in method — Local (with a password), Active Directory (with an optional directory username if it differs from the MDM username), or an identity provider (§12.3). Accounts created by SCIM (§10.10) carry their own method, have no local password, and are managed from the directory rather than here.
- Activate/deactivate accounts and reset passwords (local accounts only).
- Promoting a user to Administrator uses a four‑eyes rule once more than one admin exists — a different administrator must approve the elevation.
Deactivating a user asks for confirmation; re‑activating does not. That asymmetry is deliberate — restoring access somebody already had is not the risky direction.
Users can enable multi‑factor authentication on their own account from My Account (a time‑based code from an authenticator app). Administrators do not hold users' MFA secrets.
10.3 Roles & Permissions and Access Control
- Roles & Permissions manages role assignments.
- Access Control is the permission matrix: for every governed area and action (View / Create / Update / Delete), choose which roles are allowed, and whether a write action must go through the change‑request workflow. The matrix is grouped by layer (Administration, Model & Environment, Data Model, Source & Ingestion, Data Quality, Matching, Standardization & Reference, Governance). Role: Administrator.
10.4 Data Access (sensitive‑data clearances)
Role: Administrator. Covered in full in Chapter 11 — this tab grants per‑user clearances and shows the sensitive‑data access log.
10.5 Active Directory
Role: Administrator.
The Active Directory tab configures LDAP sign‑in for your tenant:
- Server URL (LDAPS recommended), domain or user DN pattern for direct bind, or a service account + base DN + search filter for search‑then‑bind.
- StartTLS and certificate‑verification toggles.
- A Test connection button that verifies the service bind and, optionally, a full end‑to‑end sign‑in with a test username and password.
The service‑account password is encrypted at rest and never shown again. Directory users are pre‑created by an admin (Users → sign‑in method Active Directory); their role and tenant stay under your control — only the password check goes to the directory.
10.6 Audit Trail
Role: Administrator or Steward.
The Audit Trail records who did what, when — with the resource shown by its business name (model name, user, golden key, …) rather than an internal ID. Click any row to open the full event detail: the action, the actor, timestamp and IP, and the exact recorded values (with the raw payload available for evidence). You can filter by resource type, user and date.
What the search box searches — and what it deliberately does not. Search covers username, action, resource type and resource id — the four things printed in the grid. The page states this beside the box, and the result set reports the fields it searched.
It does not search the recorded old and new values, and it does not search the resolved business name. Both omissions are on purpose:
- Searching payloads would be an unrestricted read of governed values across every record in the tenant, by anyone who can open the audit trail — including values the searcher has no clearance to see.
- The business name is resolved after the query, from the audited payload, so it is a display convenience and cannot be a search term.
So you can ask the audit trail "what did Nomsa change last Tuesday?" or "every DELETE on a golden record". You cannot ask it "find the audit entries mentioning this customer's ID number". For that, start from the record and read its history.
Audit entries cannot be edited or deleted — the row menu on this page has no delete for that reason (§2.10).
10.7 Deployment Models
Role: varies; model management typically Developer/Admin.
The Deployment Models tab manages models: create, clone, import/export (XML), archive, switch the active model, and govern each model's behaviour:
- Review bypass on/off per model (§7.3) — a row with the bypass on carries an amber Bypass ON chip with a Turn off action; a governed row shows only a quiet shield icon, which turns the bypass on.
- Strict pipeline mode on/off per model (§7.5).
- The MDM pattern the model follows — registry, consolidation, coexistence or transaction — which is the statement of how far the hub is meant to be authoritative.
- The four hierarchy tier names (§3.7) and the model's golden‑record threshold.
- Target databases are provisioned automatically when a model is created or its target changes (§2.3).
Check‑in review (change‑sets). Checking a model out takes a snapshot of its design. When you check in, the platform diffs what you changed against that snapshot and shows you the change‑set — every added, altered and removed object. You can accept or reject each change individually; rejected changes are reverted together, in one transaction, in the correct dependency order. This is what makes a checkout safe to abandon: you are never forced to keep a change you made by accident.
10.8 RDM Serving
Role: Administrator to manage; Steward and Data Owner read‑only. Manages the read‑only reference‑data serving API and its service tokens (see Chapter 9).
10.9 Health & housekeeping
The platform provides per‑hub health diagnostics (surfaced at sign‑in and in the About dialog) and housekeeping tools that scan for orphaned metadata and physical residue and clean exactly what you select, with cascade safety and a step‑by‑step report.
10.10 SCIM Provisioning
Role: Administrator.
SCIM Provisioning lets your identity platform create, update and deactivate MDM Studio accounts automatically, so joiners and leavers are handled where they are already handled.
- Issue a bearer token for your identity provider. Tokens are tenant‑scoped, stored hashed (you see the value once, at creation), and revocable at any time from this page.
- Provisioned users are created with a SCIM sign‑in method and no local password — they authenticate through your provider.
- A leaver removed in your directory is deactivated, not deleted, so their audit history and their name on past change requests survive.
Provisioning does not decide what a user may do. Roles and permissions stay with you (§10.3).
10.11 Active Sessions
Role: Administrator.
Active Sessions shows who is signed in right now — user, role, tenant, when they signed in and from where. Two actions are available on a session: send a message to that user, and end the session. Use the first before the second when you are about to restart something.
10.12 Security Posture
Role: Administrator.
Security Posture is a self‑assessment of how this deployment is configured, mapped to the control families of SOC 2 and ISO 27001. It reads the platform's own settings — encryption, sign‑in methods, session policy, masking enforcement, audit coverage — scores each control, and produces a downloadable assurance report you can hand to an auditor or a customer's security team.
Read the report for what it is. It is evidence of how this platform is configured, produced by the platform itself. It is one input to an audit, not a certification, and it says nothing about the controls around the platform — your network, your directory, your backups off‑platform.
10.13 Secrets Vault
Role: Administrator.
Every credential the platform holds — source connection passwords, the directory service account, integration keys — is stored in the Secrets Vault, encrypted with envelope encryption: each secret has its own key, and those keys are themselves encrypted by a master key.
The page shows what is stored (never the values) and supports a one‑click master‑key rotation, which re‑wraps every secret under a new master key without any secret being re‑entered.
A configuration caution worth knowing. If the vault's master key is not configured, or is configured incorrectly, the vault is simply not in use — the platform falls back to its legacy encryption and does not raise an alarm. Check this page after any change to deployment configuration: it is the only place that will tell you.
10.14 Performance and Scale Readiness
Role: Administrator or Data Owner.
- Performance benchmarks the match engine on your own hardware and data, and keeps a history so you can see the effect of a change — new blocking keys, a bigger machine, set‑based candidate generation turned on.
- Scale Readiness assesses whether the deployment is configured for the volume you intend to run: indexing, caps, pushdown settings, scheduler ownership, queue configuration. It produces a certification record of what was checked and what it found.
Both report on your deployment. The platform publishes no throughput figures of its own, and you should not assume any — run the benchmark and use your own numbers.
10.15 Backup & Restore
Role: Administrator.
Backup & Restore protects the thing that is hardest to rebuild: your modelling backbone. A backup archives models, domains, entities, attributes, mapping sets, managed values, attribute sources, survivorship configuration, glossary and DQ rulesets and rules into one versioned archive.
- Backups can run on a schedule, with a retention policy, and optionally be copied to object storage off the platform.
- Restore creates a new working model. It does not overwrite the model you are in. Everything is restored with fresh identifiers and its internal references remapped, so a restored model sits alongside the original and can be compared with it before you switch.
This is a configuration backup, not a data backup. It does not contain your staged rows, your golden records or your audit trail. Those live in the hub database and are your database platform's responsibility — keep taking database backups.
To move a model's configuration to another environment rather than restore it here, see §10.20 Promotion — the same archive format, with a sealed channel, a dry run and a recorded outcome on both hubs.
10.16 Attachment Storage
Role: Administrator.
Attachment Storage is the hub's file store, and this page is where you work with it: every file that has been uploaded, how much space it is using, what still points at it, and what can safely go.
What is in here. Files reach the store from two places — a file attached to a direct message or a discussion thread, and an image used as a dashboard's background. They are held per tenant and de-duplicated by content, so uploading a file the hub already holds re-uses the one it has instead of storing a second copy.
The four figures are the filters. Files, total size, orphans and orphan size each count a set, and clicking one shows you that set — the size figures largest-first, which is usually why you opened the page.
- Open shows the file: images and PDFs render in place, and anything else offers itself for download.
- Upload adds one or more files. If the store already holds a file with identical contents, the page says so rather than claiming an upload that did not happen.
- Rename changes the name people read. The contents, the checksum and the download address are untouched — a rename is not a re-upload, and nothing that points at the file is affected.
- Delete removes one file. If anything still points at it, the confirmation says how many and names them, and the button changes to Delete anyway — a different act, differently named. Afterwards the page reports how many references were broken.
- Used by lists every place that points at the file, each with a link that takes you there.
What "orphan" means. A file is an orphan when nothing points at it — no message attachment, and no dashboard background. It is not about whether some owning record still exists; a stored file has no owner, only referrers. Clean orphans deletes exactly that set and leaves everything still in use alone.
⚠ The list on this page and the cleaner are the same question, asked once. That matters, because a dashboard's background image is not an attachment row anywhere — it is an address held inside the dashboard's own layout. A cleaner that looked only at message attachments would call every dashboard background an orphan and delete it, and the picture would then fail to a blank space rather than an error. If you add a new kind of place that can point at a stored file, it has to be taught to this page, or the cleaner will not know about it either.
10.17 RDM Integration
Role: Admin, Steward or Data Owner.
RDM Integration manages the downstream consumers of your reference data (§9):
- Subscriptions — which systems want to hear about which reference domains.
- Signed webhook push — when a release is published, subscribers are called with an HMAC signature they can verify, so a subscriber can prove the notification came from you.
- Service accounts — the identities downstream systems use, with revocable tokens.
Push complements polling; it does not replace it. A subscriber that misses a webhook still finds the new release on its next checksum poll.
10.18 Catalog Integration
Role: Admin, Steward or Data Owner.
Catalog Integration connects MDM Studio to an external enterprise catalog. It is not the Catalog tab in Modelling Studio (§3.8) — that one is internal.
- Publish out. Domains, entities, attributes and lineage are published in an Apache‑Atlas‑compatible shape. Providers are a built‑in catalog (the default, and free), Apache Atlas, Microsoft Purview and Collibra.
- Pull in. A scheduled job pulls metadata from an HTTP endpoint and imports the delta. There is a dry‑run preview before anything is written, populated values are never overwritten by an import, auto‑accepted changes are audited, and a health digest reports what each run did.
10.19 Licensing and editions
Role: Administrator. Open Licensing from the profile menu — it is not on the tab strip.
MDM Studio comes in four editions. Enterprise+ is an Enterprise licence with a higher record band and a commercial support wrapper — in the product it is the Enterprise edition, so the table below lists the three the platform itself distinguishes:
| Edition | Domains | Users | Hub engines | Notes |
|---|---|---|---|---|
| Community | 3 | 10 | PostgreSQL, MySQL | Free. |
| Professional | unlimited | 50 | + SQL Server | Priced on record band plus seats. |
| Enterprise | unlimited | unlimited | + Oracle | Priced on record band. Enterprise+ above 5 million records. |
A licence is a signed key you install on this page; the page shows the edition, its limits, its features and its expiry.
What is actually enforced. Two limits are enforced, always, at the moment you create the object: the number of domains (Community only) and the number of users. Exceeding either is refused with a clear licence message.
Your mastered-record band is metered and shown, never enforced — the platform will not refuse to master a record because a band is full. Nothing else is enforced by default. Feature entitlement is checked only if your deployment has switched enforcement on, and today the only two screens behind that gate are Source DQ and Source Comparison (§6.8, §6.9). The hub engine entitlement is not checked at all — a Community deployment will happily run on SQL Server or Oracle. Treat the engine and feature columns as a statement of your commercial agreement, not as a technical control, and do not rely on the platform to keep you inside it.
The product fails open. No licence, an unreadable licence or an expired licence puts the platform into Community limits — it never locks you out of your own data.
Where a deployment pins its licence in configuration, install and remove are disabled on this page and the platform says so.
11. Sensitive data protection
MDM Studio can classify data and control who sees it, down to individual fields and records. This chapter brings the whole feature together.
These controls are compliance‑enabling: classification, clearances, field masking, record restriction and an access log give you the mechanisms a privacy regime such as POPIA or GDPR expects you to have. Whether your organisation is compliant depends on how you configure and operate them, and on a great deal that sits outside this platform. The product provides the controls; it cannot provide the compliance.
10.20 Promotion
Role: Administrator.
Promotion moves a model's configuration from one environment to another — development to UAT, UAT to production — over a sealed channel, with a dry run before anything is written and a way back after it is. It is the same engine as Backup & Restore (§10.15): a promotion archive is a configuration archive, extended with the mastering and reporting configuration, so the two file formats are interchangeable.
The channel. The receiving hub issues a scoped token and shows it once; the sending hub seals it in its secrets vault. The receiving hub keeps only a hash of it. A token is never displayed again anywhere — a lost one is re-issued, never recovered.
- The mark is the gate. A model can only be sent once somebody has marked it a promotion candidate. That mark is the record of a decision that this model is ready to leave the environment, and it appears on the model row in Deployment Models.
- Direction is derived, not typed. Each environment is registered as development, UAT or production, and the two roles decide whether a promotion is uphill (toward production), downhill or lateral. Uphill carries the heaviest guards: a typed confirmation naming the environment, and a reason recorded verbatim in both hubs' logs.
- Always dry-run first. The dry run is where the receiving hub answers every question that can refuse a promotion — and the answers come back before a single row is written. The Promote button is not offered until it has run.
- Nothing lands on an open edit. If somebody has the target model checked out on the receiving hub, the promotion is refused and names who holds it. A promotion replaces a whole model configuration; it will not overwrite work in progress.
- A copy is taken before the first change. The receiving hub archives everything it currently holds for that model before it changes anything, and that restore point is listed under History. Restoring one takes its own restore point first, so going back is itself reversible.
What travels, and what deliberately does not. The model backbone always travels — domains, entities, attributes, mappings, survivorship, glossary and DQ rules. Extended mastering configuration (standardization, match rules, hierarchies, reference domains and their values and releases, display settings) and analytics (reports, saved views, dashboards) are opt-in and on by default. Several things are excluded by design, and the checklist lists each one with the reason rather than leaving a gap.
- Records never travel. Staged rows, golden records, cross-references, exceptions and run history stay where they are. This moves configuration, not data.
- Credentials never travel. Connection passwords, API tokens and vault contents are excluded by name.
- Environment identity never travels. Tenant records, branding and the environment registry itself are the container, not the cargo.
- Each person's own settings stay put. Display preferences travel at model level only; a promotion never replaces an individual's saved settings on the receiving hub.
- Sensitivity classifications DO travel, with the model backbone, and are not stripped. A promoted model is never less governed than the one it came from.
Two hubs on different builds. Before writing anything, the receiving hub compares the columns it has against the columns the archive carries. If the archive holds a column this hub lacks and that column carries values, the promotion is refused and names the table, the column and both build numbers — because the alternative is worse than a failure: the rows would land, the counts would match, and the field values would be silently gone. A column that is empty everywhere is reported and allowed.
Everything is recorded, on both sides. Each hub logs its own half of every attempt — sent, landed, refused and failed — naming the actor, the direction, both build numbers and the archive checksum. Refusals are logged too: a hub being probed with bad tokens leaves a record, not a silence.
If the two hubs cannot reach each other, the archive can be downloaded and carried by hand. The checksum travels inside the file, so the receiving hub's checks are identical either way.
Reference data and the rest of the extended configuration. Reference domains and their values, aliases, standards, set policies and published releases travel with the model that owns them. So do value hierarchies and crosswalks — but only the ones that are wholly this model's, and the reason is worth reading before you plan a promotion around them. A hierarchy whose levels reach into another model's domain, or a crosswalk with one end in another model's domain, belongs to two models at once; promoting it would rewrite the other model's configuration on the receiving hub, without anyone working on that model being asked. Those are left where they are.
What is left behind is named, every time. The dry run lists each hierarchy and crosswalk it is not taking, by name, with the reason it is not taking it, and the same sentence is written into both hubs' logs when the promotion lands. This is a notice, not a refusal — nothing has gone wrong, and everything else in the selection lands normally. It is there because the alternative is worse: a promotion that reports success while something you expected to arrive quietly did not, and no record anywhere of the difference. If you need one of these to travel, give it a home in a single model first.
Two more kinds of configuration travel. Fix patterns that somebody wrote by hand travel with the model; the pattern library the product ships does not, because the receiving hub already has its own and re-creates it on every start. Report sources that belong to a model travel; sources describing the receiving hub's own stage tables are that hub's own business and stay put.
One table still cannot travel, and the checklist says which and why rather than leaving a gap. Hierarchy presets are shaped by the hub they were built on and have no link to a model at all — there is nothing to promote them as.
11.1 Classification levels
Every attribute can carry a sensitivity level:
| Level | Name | Meaning |
|---|---|---|
| 0 | Public | No restriction. |
| 1 | Internal | Visible to internal users by default. |
| 2 | Confidential | Restricted; needs clearance. |
| 3 | Restricted | Highly restricted; needs clearance, and can hide whole records. |
Levels cascade down the model: a domain default applies to everything under it, an entity can raise it, and an attribute can raise it further. The effective level is the highest in that chain. Legacy PII/Sensitive flags map automatically (Sensitive → Confidential, PII → Internal) when no explicit level is set.
Set levels in the Data Dictionary (per attribute) and on the domain and entity forms (defaults) — or let the Model Builder's AI propose them at build time (§3.2).
11.2 Clearances (who may see what)
Role: Administrator, in Administration → Data Access.
A clearance grant gives a user permission to see data up to a chosen level, for a chosen scope:
- Domain — everything under a domain.
- Entity — a specific entity.
- Record — an individual golden record.
Grants can carry an expiry. A user's effective clearance for any value is the highest of the default clearance and any grants covering it. Tenant administrators see everything by default (configurable). Users without a grant see up to the deployment's default clearance (Internal by default).
11.3 What users see
- Field masking — values above a user's clearance are shown masked (e.g.
••••, or••••1234keeping the last four for contact/identifier fields), with a lock indicator. Masking applies everywhere values are served: golden records, the Record Inspector, Report Studio (grids, aggregates and exports), exceptions, match review and profiling. - Restricted records — a golden record tagged Restricted is hidden entirely from users who aren't cleared for it: it doesn't appear in lists and returns "not found" on direct access, so even its existence stays private. Grant a specific viewer access from the record, or in Data Access.
- Before publishing — records can be restricted before they are ever published (§5.2): the restriction is held on a hidden stub (status Unpublished) and enforced identically, surviving survive and publish cycles.
11.4 Access audit
The Data Access page includes an Access log showing clearance grants and revocations, record restrictions, and sensitive‑view events (when a cleared user was actually served classified values) — so you can answer "who has seen this data."
11.5 Access groups (row‑level access)
Role: Administrator, in Administration → Access Groups.
Clearances decide which values (columns) a person may see. Access groups decide which records (rows) they may see at all — row‑level security that composes with clearance. A group has members and a set of rules, each scoping to one of:
- Domain — membership of a governed domain.
- Hierarchy tier — a tier level equals a value (e.g. Tier 1 = CORPORATE).
- Value hierarchy — a level of a value hierarchy equals a value (e.g. Province L3 = Johannesburg).
- Tree node / subtree — a node in a managed hierarchy, optionally its whole subtree.
- Column value — a golden attribute equals a value (equality, never a pattern).
How rules combine. Within one group, different kinds AND together and the same kind ORs; across a user's groups, the grants union. There are no negative rules — you narrow by not granting. A rule that cannot be resolved is refused, never dropped (dropping it would silently widen access), and an empty resolved set means "no rows", never "all rows".
Arming. Access groups are opt‑in per domain: a domain does nothing until an administrator arms it. Roll out in monitor mode first, then strict. Administrators are outside the scope unless a domain sets admin‑scoped, so they are not accidentally locked out of the data they govern.
What a scoped user sees. One scope rule is applied by every serving surface, so every count and calculation aligns to the caller's access groups — the golden grid, the exception queue, quality scores, By‑tier and by‑hierarchy counts, dashboards and campaigns. Locked filter chips name the groups doing the filtering; a governed domain the user holds no group for is not listed (its absence is stated); a refusal is a named answer, never an empty grid; and a surface a user's role cannot use is not shown at all. Before you arm a domain or save a rule, use the preview — it reports "what would this grant" and "what does this person see" as visible‑versus‑total counts, so blast radius is known before enforcement.
11.5 Configuration
Enforcement is controlled by deployment settings: an on/off switch (masking is on by default), the default clearance level, and whether tenant admins implicitly see everything. Ask your platform administrator to change these.
Two of those defaults are worth stating plainly, because they surprise people:
- Tenant administrators and platform superusers see everything unmasked, without any grant. That is the shipped default. It can be turned off per deployment, and should be if your policy says an administrator is not automatically cleared.
- Pipelines and jobs are not masked. Matching and survivorship have to compare real values to work, so the engines read unmasked data. Masking is applied when data is served to a person — grids, record views, reports, exports, APIs. This is a deliberate boundary, not a gap: what it means is that clearance controls who can see a value, not which values the platform processes.
11.6 Masking policy — the shape of a redaction
A clearance decides whether a value is masked. A masking policy decides what the mask looks like. The two are separate on purpose, and only the first can hide a value.
Masking policy is set on a glossary term (§3.5), which is what makes it manageable: you set it once on National Identity Number and every attribute governed by that term inherits it. The vocabulary is closed — these five and no others:
| Policy | What a masked value looks like |
|---|---|
| FULL | Nothing survives. The default, and the fallback. |
| FIRST_INITIAL | The first character only — N••••• — enough to disambiguate two records in a list. |
| YEAR_ONLY | For a date, the year alone — 1987 — enough for an age band, not enough to identify. |
| LAST4 | The last four characters — ••••1234 — the familiar shape for account and card numbers. |
| EMAIL_DOMAIN | The domain only — •••@example.co.za — enough to tell which organisation, not who. |
Three rules govern how they behave:
- A policy never decides whether to mask. There is deliberately no "none", "clear" or "off" in the list. If you want somebody to see a value, grant them a clearance (§11.2); a masking policy cannot do it for you.
- Where terms disagree, the most restrictive wins — and the platform reports which policy it applied, so a surprising redaction can be traced to the term that caused it.
- An unrecognised policy fails closed to FULL. A typo makes a field more hidden, never less.
A term's policy overrides the default shape for its classification level, and it can reveal more than the default would — LAST4 on an account number where the level alone would have hidden everything. That is a real decision with real consequences, which is why it is written on the governed term where it can be reviewed, rather than typed into a field on a screen.
Every response that carries a redacted value also carries which policy produced it, so grids, exports and support investigations can all answer "why does this look like that?".
12. Platform & operations reference
This chapter summarises the operational facts most useful to administrators. Full deployment steps live in the deployment pack that ships with each release.
12.1 Engines
One canonical model runs natively on SQL Server, PostgreSQL, MySQL and Oracle through a per‑statement dialect translation layer. The engine is a per‑session choice at sign‑in; connections and source systems are shared across engines automatically. Model target databases and their pipeline schemas are auto‑provisioned on SQL Server, PostgreSQL and MySQL (§2.3).
12.2 Multi‑tenancy
Tenants are fully isolated: data, users and models are scoped to a tenant, with per‑tenant branding, account‑filtered sign‑in and full suspension. Deployment can be in your own infrastructure or as an isolated tenant on a managed hub.
12.3 Sign‑in security
- Local sign‑in — passwords hashed in MDM Studio.
- Active Directory / LDAP, configured per tenant, by direct bind or search‑then‑bind (§10.5). An empty‑password bind is refused outright, because some directories treat it as a successful anonymous bind.
- OIDC — authorisation‑code flow, with accounts provisioned on first sign‑in and directory groups mapped to MDM Studio roles.
- SAML 2.0 — HTTP‑POST binding, RSA‑SHA256 signature verification, audience and conditions enforced, and a replay guard.
- SCIM 2.0 for provisioning rather than sign‑in (§10.10).
- Multi‑factor authentication — self‑service, time‑based codes, enabled by the user in My Account.
- Idle sign‑out with a warning dialog (§2.1); the timeout is a deployment setting.
- A lockout‑proof recovery administrator: deployments can guarantee a working admin account on every start, so the platform can always be recovered.
OIDC and SAML are inactive until a deployment configures them. If your sign‑in page offers only local and directory sign‑in, that is why. Validate either against your own identity provider's metadata as part of rollout — the platform implements the profile, but only your provider can confirm the pairing.
12.4 Versioning & About
MDM Studio ships as one versioned release covering the web studio and the backend together, using SemVer — the current release is v1.2.0. The About dialog shows the running version, the active hub engine and the platform capabilities; a Copy info button copies these details for support. If the studio and backend versions ever differ, About warns you to deploy the matching pair.
About also shows two build numbers, one for the studio and one for the backend. These are internal component counters for support diagnostics — they identify exactly which build you are running when you raise a ticket. They are not the release. When you record what you are running, in a change record or a conversation with your vendor, the answer is v1.2.0; quote the build numbers only when support asks for them.
12.5 Defaults that change what you see
Several capabilities are configured per deployment, and their defaults explain many "why is this not happening?" questions. Ask your administrator which of these apply to you.
| Behaviour | Default | What that means for you |
|---|---|---|
| Catalog index refresh | Manual only | The catalog updates when somebody presses Reindex. A drain interval is optional; without it the index is exactly as current as the last reindex (§3.8). |
| Catalog vector search | Off | Search is the deterministic keyword ranking only; the studio says the lane is off. |
| Catalog embedding model | Local | Lexical matching in‑process, with no data leaving the deployment. A learned model requires explicit configuration. |
| Feature entitlement enforcement | Off | Editions do not gate features, apart from domain and user counts (§10.19). |
| Durable job queue | Off | Scheduled and triggered runs are fire‑and‑forget. A run interrupted by a restart is lost, not retried — check the run history after any restart. |
| Scheduler | On | Must be turned off on every host but one in a multi‑host deployment, or each host fires every schedule. |
| Tenant isolation guard | Monitor | A query that slipped past tenant scoping is logged, not blocked. Fail‑closed mode is available and should be adopted after a staging soak. |
| Masking | On | Values above your clearance are masked when served (§11.3) — but administrators are implicitly cleared by default (§11.5). |
| Secrets vault | Configured or silently absent | A malformed vault key is treated as no vault at all, with no error. Confirm on the Secrets Vault page (§10.13). |
| Multi‑process clustering | Off | One process. Turn it on to use more than one core. |
12.6 What is AI, and what only sounds like it
MDM Studio contains a great many advisors. Almost all of them are deterministic: they read your profiled data and your model, apply published rules, and explain their reasoning. The same input always produces the same advice, and no data leaves your deployment. The survivorship, match, DQ‑rule, standardization, hierarchy‑tier, crosswalk, trust, golden‑record, domain‑doctor, placeholder and policy advisors are all in this group.
A small, named set of surfaces uses a large language model, and only when your deployment has configured a key: catalog description and adjudication suggestions, the catalog's Ask box, the optional remote embedding model, the annotation beside the placeholder advisor's arithmetic, the annotation beside the policy advisor's arithmetic, AI naming in the Model Builder, the two hierarchy‑tier advisors, and the Help assistant.
What is sent, when those are used: asset names and types, numeric statistics and pass rates, candidate vocabulary, descriptions your stewards wrote, and the names and definitions of glossary terms.
What is never sent: any data value, any row or record, any golden record payload, any identifier, and any minimum, maximum or frequency value. Nor is the question you typed — a free‑text question is exactly where a real data value tends to appear ("who owns the record for 8801235459081?"), so questions are interpreted inside the platform and only the retrieved metadata is sent onward.
With no model configured, nothing breaks. Every AI surface has a deterministic fallback: naming falls back to a rule‑based expander, tier advisors fall back to a curated industry library, the Help assistant falls back to ranked sections of this guide, catalog descriptions return nothing and say so, and the catalog agent answers from the deterministic layer. Each result reports which engine produced it, so you always know whether you are reading arithmetic or a suggestion.
One phrase to be careful with. Catalog search in a stock deployment is lexical, not semantic — it matches shared words and word‑shapes, and the platform reports it as non‑semantic in its own status. Nothing in the product infers meaning from your data unless a learned model has been deliberately configured.
12.7 Before you upgrade — two behaviour changes
Role: Administrator. Read this before installing this release over a hub whose pipelines are already running.
Two changes in this release alter what running pipelines produce, and each ships with a report that names exactly what will change on your hub. Run both reports before the first post‑upgrade pipeline run.
- Transforms now execute. Previously, most per‑field transforms (UPPER, LOWER, TRIM, EXPRESSION) were accepted by the designer but silently ignored at run time. From this release the generated extract SQL applies them, so a mapping that carries a transform can land different values on its next run. Nothing already landed is rewritten — only future runs change. The transform‑impact report (
GET /mappings/transform-impact) names every mapping whose output moves. CONCAT and LOOKUP remain non‑executing and are now refused on save with the reason (§4.2). - Approved specification versions govern runs. Approving a mapping set freezes an immutable version, and the runner executes the approved version, not the draft. A set with no approved version keeps running its draft exactly as before — so on most hubs, nothing changes at install time. The spec‑impact report (
GET /mappings/spec-impact) names the sets whose drafts have diverged from their approved version; on those sets, the approved version is what runs until you approve again. In that report, drifted: null means not measured, not "no drift" (§4.2).
Neither change requires action to stay safe — the reports exist so that the first changed value is a decision, not a surprise.
Appendix A — Roles & permissions
Default role capability (deployments can customise this in Access Control). "•" indicates typical access to that layer's actions.
| Layer | Viewer | Analyst | Steward | Data Owner | Developer | Admin |
|---|---|---|---|---|---|---|
| View data & reports; Report Center, GR Explorer | • | • | • | • | • | • |
| My Work, Messages, My Account | • | • | • | • | • | • |
| Data model (domains, entities, dictionary, hierarchies) | • | • | • | • | ||
| As‑Is Assessment | • | • | • | • | ||
| Catalog — search and read | • | • | • | • | • | • |
| Catalog — edit a facet (governed) | • | • | • | • | ||
| Sources & ingestion | • | • | • | • | ||
| Standardisation & reference | • | • | • | • | ||
| Matching & survivorship | • | • | • | • | ||
| Author Records | • | • | • | |||
| Write‑Back, Real‑time API | • | • | • | • | ||
| Data quality, exceptions, rule suspension | • | • | • | • | • | |
| Source DQ, Source Comparison, Placeholder Advisor | • | • | • | • | ||
| Governance (submit/approve), Approval Workflows | • | • | • | • | ||
| Report Studio (build) | • | • | • | • | • | |
| RDM Serving, RDM Integration, Catalog Integration | • | • | • | |||
| Audit Trail | • | • | ||||
| Performance, Scale Readiness | • | • | ||||
| User management, Roles & Permissions | • | • | ||||
| Access Control, Data Access, Active Directory, SCIM | • | |||||
| Security Posture, Secrets Vault, Active Sessions | • | |||||
| Backup & Restore, Attachment Storage, Licensing | • | |||||
| Tenants | • |
Platform superuser (separate from role) is required for cross‑tenant administration (the Tenants page).
Two shaping rules sit on top of this table: a tab your role cannot open is hidden, not greyed; and for stewards, data owners and analysts the modelling, integration and administration hubs sit inside the collapsed Engineering disclosure (§1.3).
Appendix B — Glossary
- Attribute — a field of an entity, defined in the Data Dictionary.
- Catalog facet — a piece of curated metadata on a catalog asset (owner, steward, consulted, informed, classification, tags). Set‑valued: saving replaces the whole set. Editing one is a governed change (§3.8).
- Change request — a governed change awaiting review/approval.
- Change‑set — the diff between a model as you checked it out and as you are checking it in; each change can be accepted or rejected individually (§10.7).
- Checkout — an exclusive edit lock on a model.
- Clearance — a user's permission to see sensitive data up to a level, per scope.
- Confidence — a golden record's score computed from real match evidence (pair scores and field agreement).
- Cross‑reference (xref) — the link from a golden record to a contributing source record.
- Domain — a subject area grouping entities.
- Engine / hub — the database platform storing the master‑data hub.
- Entity — a master‑data object producing golden records.
- Exception — a data‑quality rule failure to be remediated.
- Golden record — the single trusted mastered record.
- Hierarchy (managed) — a maintained tree over master data, with versions, diffs and integrity checks (§3.7).
- External asset — a curated record of something outside MDM Studio that matters to your lineage: a warehouse, a table, a file, a report. Addressed
ac:x/{system}[/{name}], held globally (no model), linked to your own assets by CONSUMES / FEEDS / DESCRIBES, audited rather than review‑gated, and retired to a tombstone rather than deleted (§3.9). - Hierarchy tier — one of four organising levels bound to an attribute and stamped through every pipeline stage (§3.7). Not the same thing as a managed hierarchy. Where the bound attribute is standardized, the tier is stamped from the cleansed value (§3.7).
- Value hierarchy — a declared containment chain between governed code sets (Gauteng contains Johannesburg), with levels that are reference domains and per‑entity bindings separate from tier bindings (§3.10).
- Fix pairing — the governed registry entry pairing a DQ rule's operator with a candidate standardization transform: FIXES, ENABLES or NONE, each with its reason (§6.4).
- Master standardized value — the per‑attribute flag that publishes the cleansed value into the golden record. Defaults on at the first rule attached to an undecided, eligible attribute; a recorded human decision is never overturned (§5.1).
- Label attribute — the attribute whose value names a hierarchy's nodes (the category name, not its code). Configured when a hierarchy is created from a relationship, editable on every surface that renders the tree (§3.7).
- Masking policy — the shape a redaction takes: FULL, FIRST_INITIAL, YEAR_ONLY, LAST4 or EMAIL_DOMAIN. It never decides whether a value is masked (§11.6).
- Model — a complete versioned snapshot of the design.
- Pipeline stage — one step of Extract → Land → Stage → Standardize → Match → Survive → Publish (ext / LND / STG / STD / MTC / SRV / GRD).
- Proposal — a suggestion produced by the catalog or an advisor, held pending until a person explicitly accepts it; a rejection is remembered so it does not resurface (§3.8, §6.7, §7.1).
- RACI snapshot — the record, written onto a governed request at the moment it is raised, of who was Accountable, Responsible, Consulted and Informed then. An audit fact, not a live view (§7.7).
- Record Inspector — the overlay showing a record's cluster, related records, exceptions and lineage in one place.
- Reference data (RDM) — governed shared code sets with immutable releases.
- Retired (record) — a stage row soft‑removed from the pipeline; retired rows never flow to later stages.
- Review bypass — a per‑model switch that suspends governance review, so governed changes apply directly instead of requiring change‑request approval. Bypass off is normal governance (§7.3).
- Sensitivity level — Public / Internal / Confidential / Restricted classification.
- Source object — a registered source table, file or endpoint with a snapshotted field inventory, scope marks and drift detection (§4.1).
- Source system / connection — an origin of data and the credentialed connection to it.
- Specification version — the immutable, numbered snapshot of a mapping set frozen at approval. The pipeline executes the approved version; the draft stays editable and runs only where no version was ever approved (§4.2).
- Strict pipeline mode — a per‑model gate making Survive/Publish require their governance prerequisites (§7.5).
- Superseded — a golden record whose source record disappeared; kept, not deleted, and revived if the record returns.
- Survivorship — the rules deciding which source value wins in a golden record.
- Suspension window — a DQ rule paused until a stated date, with a mandatory reason, returning by itself. A suspended rule is reported as not evaluated, never as passing (§6.6).
- Target kind — whether a mapping set's target is one of your model's entities (ENTITY) or a registered external asset (EXTERNAL). Exactly one, never both. An EXTERNAL‑target specification is governed and versioned like any other, and is refused by name by every pipeline (§4.2).
- Tenant — an isolated workspace.
- Tier binding — the link from a hierarchy tier to the attribute that supplies its value for a given entity (§3.7).
- Tombstone — a catalog asset whose source object has vanished. It is marked as gone rather than deleted, so its history and its human‑set facets survive (§3.8).
- URN — a catalog asset's stable address, e.g.
ac:entity/CUSTOMER/EMAILorac:crm/sales/dbo/CUST/EMAIL. It is the join key across the catalog (§3.8).
Appendix C — Keyboard & navigation
- ⌘K / Ctrl‑K — open the command palette (search and jump anywhere).
- Right‑click a grid row — the row menu for that row and that column (§2.10). Escape closes it; so do a click outside, a scroll or a resize.
- Backspace or the ← Back button (hub tab strip) — step back through the pages you visited.
- Refresh (emerald button, on every page) — force a fresh load; navigation is otherwise served from a smart cache.
- Sidebar — switch hubs; collapse to icons with the chevron.
- Model switcher (header) — change the active working model.
- Profile menu — My Account, tenant info, sign out.
Appendix D — Status & indicators
- Amber "Review bypass ON" pill + amber header theme — governance review is suspended for the active model; governed changes are applying directly, unreviewed (§7.3). A plain header with no pill at all means normal governance protocols — the header speaks only when something is switched off.
- Amber "Bypass ON" chip on a model row — that model's review bypass is on; the Turn off action beside it restores governance.
- Quiet shield icon on a model row — that model is governed (bypass off); click it to suspend review.
- Lock icon on a value /
••••— the field is sensitive and above your clearance. - "Checked out by you" / "Checked out · " — model edit‑lock status.
- STAGED / WARN / RETIRED — record lifecycle statuses on stage rows (§4.4).
- ACTIVE / SUPERSEDED / MERGED / UNPUBLISHED — golden‑record statuses (§5.2).
- "Stale — config changed" — a pipeline stage whose configuration changed since its last run (§5.4).
- Delta BAU / Full‑load BAU / undecided (amber) — an Extract set's load‑expectation badge (§4.2).
- Severity dots (Critical / High / Medium / Low) — exception and quality severities.
- Suspended until \<date> — a DQ rule inside a suspension window; it will return by itself (§6.6).
- Not evaluated — a rule that produced no pass rate this run, because it is suspended or could not be evaluated. It is not a pass (§6.6).
- Proposed — an advisor or catalog suggestion awaiting an explicit accept. Nothing has been applied (§3.8, §6.6, §6.7, §7.1).
- Tombstoned — a catalog asset whose source object has gone; kept for its history and facets (§3.8).
- MANUAL — provenance on a value authored directly in the hub rather than survived from a source (§5.5).
- AUTO_MERGE / REVIEW / NO_MATCH — a match decision, in batch and in the real‑time API (§5.1, §5.7).
- Publish line — the marker on a golden grid showing how current the published set is relative to the last pipeline run.
- Population badge — how much of an entity's golden set is populated for the shown columns.
- "Not notified" on a request card — an Informed name on the RACI snapshot that matches no active user. The names are listed (§7.7).
- Greyed row‑menu item with a reason — the action is unavailable, and the text says why (§2.10).
- Record‑count pill / "N below" — on a hierarchy or tier node: how many records carry this node's own value, and how many sit in the subtree beneath it. An absent pill means not recorded, never zero (§6.2).
- "Unparented (N records)" divider — records in a record tree whose parent value resolves to nothing. A finding kept visible, not roots in disguise (§8.4).
- "tree capped at N nodes" — the serialization cap was reached; the branch below is not loaded, and is never drawn as a leaf (§8.4).
- "— not applied" on a transform — CONCAT and LOOKUP do not execute at run time; the picker and any saved row carrying them say so (§4.2).
- Diverged (specification) — the editable draft no longer matches the approved version; the approved version is what runs (§4.2).
- BREAKING on a removed source field — the field vanished from the source but a live mapping still reads it (§4.1).
- Violet dashed / amber dot‑dash edge on the ERD — declared / proposed value‑hierarchy containment; a proposed edge is click‑to‑review (§3.1, §3.10).
- "standardized for grouping — stored value is raw" — a tier bound to a standardized field that does not yet master its cleansed value; the tier groups by the cleansed form until the field is mastered (§3.7, §5.1).
- DERIVED on a rule candidate — the expression was composed from your sentence in the intent box, not from anything governed or approved (§4.3).
- rewritten in a transform preview — the value still passes its rules but the stored value changed; the quiet outcome to read before accepting (§4.3).
- SLA due / breached — on exceptions, workflow steps and My Work items.
Appendix E — Troubleshooting & FAQ
I can't see a tool / a page says my role isn't permitted. Access is role‑based and configurable. Ask an administrator to review Access Control for that area, or your role assignment in Users.
A page says "select or create a working model." The tool is model‑specific. Choose a model in the header's model switcher, or create one in Deployment Models.
A field shows ••••. It's a sensitive field above your clearance. If you need access, ask an administrator to grant you a clearance in Data Access.
A golden record I expect isn't in the list. Check the browser's status filter — the record may be Superseded (its source record disappeared; it revives if the record returns). It may also be a Restricted record you aren't cleared for; ask an administrator for a per‑record grant.
A Survive or Publish run failed with a governance message. The model has strict pipeline mode on and a prerequisite is missing (no recorded match run, or open critical exceptions). The failure panel links to exactly what's needed; an administrator can use the audited Override & run if it truly can't wait (§7.5).
The Match cluster tab says the record forms a cluster of one. That's normal for a record that matched nothing — it has no duplicates. Run the Match stage if you expected duplicates to group here.
A page seems out of date. Navigation is served from a cache for speed. Click the emerald Refresh button to force a reload (edits, model switches and sign‑ins refresh automatically).
I was signed out unexpectedly. The idle timeout signed you out after a period of inactivity. Sign back in; ask your administrator if the timeout is too short.
My changes need approval before taking effect. The action is governed by a change‑request workflow. Submit the change; an authorised approver will review it in Governance → Change Requests.
Active Directory sign‑in fails. Confirm your account's sign‑in method is Active Directory and that your administrator has enabled and tested the directory in Administration → Active Directory. If your organisation uses SAML or OIDC and neither appears on the sign‑in page, those methods have not been configured for this deployment (§12.3).
A hub or tab I was told about isn't in my sidebar. Three possibilities, in order of likelihood. If you are a steward, data owner or analyst, Modelling Studio, Integration Hub and Administration are inside the collapsed Engineering disclosure (§1.3). If it is Messages or Licensing, they open from the header and the profile menu respectively, not the sidebar. Otherwise your role cannot open that tab, and hidden is how the studio shows that — ask an administrator about Access Control (§10.3).
The catalog doesn't show a change I just made. The catalog index is rebuilt, not live: press Reindex (§3.8). Note that changes to connections, mapping sets, relationships, attribute sources and source trust scores do not even flag the catalog as stale — after that kind of work, reindex deliberately.
Catalog search isn't finding something I know is there. Search is keyword matching, not meaning: it will not find client when the asset is called customer unless the two are linked as synonyms in the glossary (§3.5). Add the synonym, and the search will find it from then on.
A facet I set has disappeared after a rename. Renaming a table or column gives it a new catalog address, and hand‑typed facets do not follow (§3.8). Re‑apply them on the renamed asset.
A rule shows no pass rate and raised no exceptions. It is suspended, or it could not be evaluated. Check the rule for a Suspended until badge and its reason (§6.6). A suspended rule is reported as not evaluated — it is not passing, and the aggregate says how many rules were skipped.
My quality score went up and I didn't fix anything. Check whether rules were suspended (§6.6) — the aggregate reports how many it did not evaluate, and that count is the thing to compare, not the percentage.
Write‑back produced a file instead of updating the source. That is the default: write‑back generates a drift‑guarded UPDATE script for a source DBA to run. DIRECT and PROC modes exist and must be chosen deliberately (§5.3).
Write‑Back says it isn't available. Write‑back is not supported when the master‑data hub runs on Oracle (§5.3).
Source DQ or Source Comparison is missing. Those two screens are the only ones gated by edition. Check Licensing (§10.19) or ask an administrator.
The real‑time API returned a golden record that's out of date. The API serves the last published master. It decides and serves; it never masters. If a source changed since the last Match → Survive → Publish run, run the pipeline (§5.7).
The real‑time API returns fewer candidates than I expected. A synchronous match loads at most 5,000 candidate rows. Add or tighten a blocking key on the ruleset so the candidate set is bounded by design (§5.1, §5.7).
A survive or publish run failed on a row cap. Runs refuse rather than truncate — that message means your data exceeded a configured cap, not that something is broken. Ask your administrator to raise the cap for this deployment (§5.1).
A scheduled job didn't run, or ran twice. If it ran several times, more than one host has the scheduler enabled — exactly one should (§12.5). If a run vanished after a restart, the durable job queue is off and interrupted runs are not retried; re‑run it and ask your administrator about enabling the queue.
The audit search box isn't finding what I search for. It searches username, action, resource type and resource id only — deliberately not the recorded values, and not the business name (§10.6). To trace a specific record, start from the record and read its history.
I deleted something and it's still there. Deletion is frequently supersession rather than removal: standardization rules are soft‑deleted and raise a change request, catalog assets are tombstoned, golden records are superseded, campaigns cannot be deleted at all, and a governed artefact is created inactive until it is approved. Check the page's status filter before assuming a failure.
Someone on my Informed list wasn't notified. The request card lists, by name, anyone on the RACI snapshot who matched no active user (§7.7). Add or activate the account and the next request will reach them; the snapshot on the existing request is a record of what was true then, and does not change.
A per‑tier scorecard or the survivorship decision matrix is empty. Tiers are not bound on that entity. Bind them in Domains & Entities or the Data Dictionary (§3.7).
The Tiers lens in By‑hierarchy is disabled. The chip states its reason: no attribute of the entity is bound to a hierarchy tier. Bind one in Hierarchies & Tiers (§3.7) and the lens lights up.
By‑tier shows a "(level unknown — recompute scores)" group. Those tier scores were captured before the level was recorded on snapshots. Press Compute scores once and they land in their proper level groups (§6.2).
Values changed after an upgrade, and I didn't touch the mappings. Transforms now execute (§12.7). Run the transform‑impact report — it names every mapping whose output moved. Nothing already landed was rewritten.
I edited a mapping set and the pipeline ignored my change. The set has an approved specification version, and the approved version is what runs (§4.2). The Versions panel will say diverged; approve the draft to make it current, or leave it — the approved behaviour is the governed one.
The Hierarchy control is missing on an entity. That entity has no navigable hierarchy source — no tier bindings, no self‑relationship, no managed hierarchy. Absence is deliberate; there is no disabled toggle (§8.4). If the data looks like a tree, watch for the measured declare‑this‑hierarchy suggestion (§3.7).
The Placeholder Advisor has nothing to say. It reads a profiling run. Run Data Profiling on the entity first (§6.1, §6.7).
Report Center, GR Explorer or Record Timeline is empty on a new deployment. They read published golden records. Complete a Match → Survive → Publish run, and — for Report Center — publish at least one report.
A page returned "no working model is active". Most tools are model‑specific and a brand‑new tenant has no model. Create and activate one in Deployment Models, then check it out before making structural changes (§2.3, §2.4).
A connection type is missing from the picker. The five cloud/SaaS connectors — Snowflake, BigQuery, Databricks, Salesforce and Teradata — ship as optional components. Ask your administrator to install the driver for that platform (§4.1).
An external asset I registered isn't in the catalog. The catalog is rebuilt, not live — press Reindex (§3.8). If it still isn't there, check the asset is ACTIVE: only active external assets are emitted, and a retired one is tombstoned by design (§3.9).
I can't create an ingestion job for a mapping set. Check the set's target kind. A set targeting an external asset is refused by every pipeline, by name and with the reason (§4.2) — that is the design, not a fault. Moving data outward is Write‑Back (§5.3).
External assets aren't being proposed from my external catalog. The import bridge ships switched off. An administrator turns it on; until then Catalog Integration pulls metadata without offering it to the register (§3.9, §10.18).
The External specifications section shows NOT MEASURED. That hub's specification version tables predate this release, so there is nothing recorded to report yet. It is not a zero, and the section distinguishes the two deliberately (§8.3).
A hub's tab bar shows groups instead of tools. The bar could not fit that hub's tools in your window, so it switched to carrying the groups with a tool row beneath (§1.3). Widening the window returns it to the flat bar on its own.
MDM Studio User Guide · Adaptive Canvas · v1.2.0. For the latest platform information, open the About dialog in the app or visit adaptivecanvas.co.za.