06 · Data platform and master data · Text-to-query over governed structured data

Governed Data Platform

Trusted data and semantic foundations that make AI and decision systems usable.

Conceptual illustration of source records connected to an organised, checked structure and a simple report display.
Conceptual illustration

At a glance

Status
In place
What this label is based on

“medallion architecture with certified Gold datasets”

The business story

Inconsistent source definitions made reconciliation and reporting difficult.

A data platform is only as good as the names, rules and lineage people can trust. The engineering here is as much governance as it is pipelines, and it is what lets agents and reports downstream use the right column and the right rule.

My contribution

  • Designed the medallion layout, naming and certification standards
  • Wrote the dictionary method and ran discovery with finance and operations subject-matter experts
  • Led data-quality diagnosis and durable remediation with subject-matter experts
  • Built the change agent and the query builder, and remediated the semantic models

The process

  1. Discovery

    Discovery sessions with finance and operations experts, and profiling of every column or measure before it was used.

  2. Mapping

    A naming standard, certification rules and the structured dictionary: join keys, rules and exclusions, metric formulas, thresholds, glossary.

  3. Build

    A change agent that turns a request into verified, human-gated commands across source, Gold and model layers, and a schema explorer with join detection.

  4. Validation

    A correction register designed to survive refresh, with the postmortem recorded, and row-volume guards in the query builder.

How it works

The governed foundation: a medallion lakehouse layout with a naming standard and certification rules, a structured data dictionary written with subject-matter experts, a source-truth discovery discipline, a refresh-durable correction register for approved data-quality remediation, a semantic-model change agent that produces verified and human-gated commands, and a browser tool that lets analysts build safe queries from a searchable dictionary.

Source data lands in Bronze, is conformed and quality-checked in Silver, and becomes certified, consistent measures in Gold that feed reporting, decisions and AI workflows. Schematic workflow, not a product screenshot.
Schematic workflow of the Governed Data Platform, not a product screenshot or a measured result.
01PROBLEM02RULES + DATA03ENGINE04AI05GATE06WORKFLOW07OUTCOMESource systemsdisagree; namesdriftNaming standard,structureddictionaryBronze, Silver,Gold pipelinesText-to-queryover thedictionarySource-truthdiscovery,change-agentverificationSemantic models,reports, agentsTrusted figuresdownstream 01PROBLEM02RULES + DATA03ENGINE04AI05GATE06WORKFLOW07OUTCOMESource systems disagree; namesdriftNaming standard, structureddictionaryBronze, Silver, Gold pipelinesText-to-query over thedictionarySource-truth discovery,change-agent verificationSemantic models, reports,agentsTrusted figures downstream
  • Deterministic stage
  • AI stage
  • Gate: human decision point
AI role AI reads the dictionary. It does not guess the join.
  1. 01 Problem

    Source systems disagree; names drift

    Source consistency needed a shared quality framework.

  2. 02 Rules + data

    Naming standard, structured dictionary

    A published table-naming reference and a dictionary with schema and join keys, rules and exclusions, metric formulas, thresholds and glossary.

  3. 03 Engine

    Bronze, Silver, Gold pipelines

    Bronze landing (named, typed), Silver conformance and quality rules, Gold certified datasets, semantic models in Direct Lake mode.

  4. 04 AI

    Text-to-query over the dictionary

    AI-generated SQL and DAX read the dictionary, so the model joins on the right key and applies the right exclusion.

  5. 05 Gate

    Source-truth discovery, change-agent verification

    Count before you select: every column or measure is profiled before it is used. A change agent turns a request into verified, human-gated commands.

  6. 06 Workflow

    Semantic models, reports, agents

    Semantic models, reports and agents are built to consume certified Gold.

  7. 07 Outcome

    Trusted figures downstream structural

    Figures people act on are designed to come from one governed layer, not whichever extract was handy.

Step by step
  1. Bronze lands source data named and typed; Silver applies conformance and quality rules; Gold holds certified datasets.
  2. The dictionary has documented sections: schema with exact join keys, business rules and exclusions, metric formulas, threshold logic and glossary. It is written with an expert discovery script.
  3. Source-truth discovery: empirically count and profile every column or measure before it is used in SQL, DAX or M.
  4. A durable correction register supports approved source-data remediation without embedding changes in a report.
  5. The semantic-model change agent turns a colleague's request into exact, verified commands across source, Gold and model layers, with refresh diagnostics, and waits for a human before running them.
  6. A schema explorer and query builder with join detection and row-volume guards lets analysts query the ERP database safely.

AI and engineering judgement

Where AI is used

Generating SQL, DAX and M against the dictionary, diagnosing refresh and freshness problems, and drafting the change commands a human then approves. When the model joins on the wrong key, the fix is a documentation gap, not a bigger model.

What remains deterministic

  • Table naming, layering and certification rules
  • Quality rules and the correction register
  • Join keys, exclusions and metric formulas in the dictionary
  • Row-volume guards and the human gate on every change command

How risk is controlled

  • Verify before you write: profiling evidence is required before a column or measure is used
  • Change commands are generated, verified and shown to a human before execution
  • Correction register designed to survive refresh, with the postmortem recorded
  • Row-volume guards in the query builder so an analyst cannot accidentally pull the whole warehouse
  • Health review and remediation of finance semantic models
Verification and controls
  • Published table-naming reference and a pipeline architecture review against the standard
  • Workspace monitoring page for refresh and freshness
  • Semantic models developed and health-checked through a modelling toolchain driven from Claude Code
  • Reusable skills for hierarchy drift diagnosis and source-truth discovery

Tech stack

Data

  • ERP and data-warehouse sources source systems with inconsistent names and rules
  • Microsoft Fabric and OneLake the Bronze, Silver and Gold layout with naming and certification rules
  • Dataflows and pipelines Bronze landing, then Silver conformance and quality rules
  • PySpark, T-SQL, M transformations written after each column is profiled

Application

  • Power BI, DAX and Direct Lake semantic models built to consume certified Gold
  • Tabular modelling tools semantic-model health checks and fixes

AI at runtime

  • Text-to-query over the dictionary AI-generated SQL and DAX read the dictionary, so the model joins on the right key
  • Semantic-model change agent drafts the change commands a person approves

AI-assisted development (not at runtime)

  • Claude Code drives the modelling toolchain that develops and health-checks semantic models
  • Reusable skills shared skill definitions

Adoption and outcomes

A governed data foundation supports consistent definitions, quality review and reusable information. Confidential operational details are withheld. structural

  • Outcome

    Bronze to Gold

    medallion architecture with certified Gold datasets

    realised

  • Engineering

    Shared definitions

    a governed dictionary supports consistent interpretation

    structural

  • Control

    Human-gated

    every agent-generated change command is verified before it runs

    structural

Status In place

What I learned

Encode business rules in a dictionary the model reads, not in prompts someone remembers. When the AI joins on the wrong key, the fix is a documentation gap, not a bigger model.

Related work