06 · Data platform and master data · Text-to-query over governed structured data
Governed Data Platform
Trusted data and semantic foundations that make AI and decision systems usable.
At a glance
- Status
- In place
What this label is based on
“medallion architecture with certified Gold datasets”
The business story
Inconsistent source definitions made reconciliation and reporting difficult.
A data platform is only as good as the names, rules and lineage people can trust. The engineering here is as much governance as it is pipelines, and it is what lets agents and reports downstream use the right column and the right rule.
My contribution
- Designed the medallion layout, naming and certification standards
- Wrote the dictionary method and ran discovery with finance and operations subject-matter experts
- Led data-quality diagnosis and durable remediation with subject-matter experts
- Built the change agent and the query builder, and remediated the semantic models
The process
-
Discovery
Discovery sessions with finance and operations experts, and profiling of every column or measure before it was used.
-
Mapping
A naming standard, certification rules and the structured dictionary: join keys, rules and exclusions, metric formulas, thresholds, glossary.
-
Build
A change agent that turns a request into verified, human-gated commands across source, Gold and model layers, and a schema explorer with join detection.
-
Validation
A correction register designed to survive refresh, with the postmortem recorded, and row-volume guards in the query builder.
How it works
The governed foundation: a medallion lakehouse layout with a naming standard and certification rules, a structured data dictionary written with subject-matter experts, a source-truth discovery discipline, a refresh-durable correction register for approved data-quality remediation, a semantic-model change agent that produces verified and human-gated commands, and a browser tool that lets analysts build safe queries from a searchable dictionary.
- Deterministic stage
- AI stage
- Gate: human decision point
-
01 Problem
Source systems disagree; names drift
Source consistency needed a shared quality framework.
-
02 Rules + data
Naming standard, structured dictionary
A published table-naming reference and a dictionary with schema and join keys, rules and exclusions, metric formulas, thresholds and glossary.
-
03 Engine
Bronze, Silver, Gold pipelines
Bronze landing (named, typed), Silver conformance and quality rules, Gold certified datasets, semantic models in Direct Lake mode.
-
04 AI
Text-to-query over the dictionary
AI-generated SQL and DAX read the dictionary, so the model joins on the right key and applies the right exclusion.
-
05 Gate
Source-truth discovery, change-agent verification
Count before you select: every column or measure is profiled before it is used. A change agent turns a request into verified, human-gated commands.
-
06 Workflow
Semantic models, reports, agents
Semantic models, reports and agents are built to consume certified Gold.
-
07 Outcome
Trusted figures downstream structural
Figures people act on are designed to come from one governed layer, not whichever extract was handy.
Step by step
- Bronze lands source data named and typed; Silver applies conformance and quality rules; Gold holds certified datasets.
- The dictionary has documented sections: schema with exact join keys, business rules and exclusions, metric formulas, threshold logic and glossary. It is written with an expert discovery script.
- Source-truth discovery: empirically count and profile every column or measure before it is used in SQL, DAX or M.
- A durable correction register supports approved source-data remediation without embedding changes in a report.
- The semantic-model change agent turns a colleague's request into exact, verified commands across source, Gold and model layers, with refresh diagnostics, and waits for a human before running them.
- A schema explorer and query builder with join detection and row-volume guards lets analysts query the ERP database safely.
AI and engineering judgement
Where AI is used
Generating SQL, DAX and M against the dictionary, diagnosing refresh and freshness problems, and drafting the change commands a human then approves. When the model joins on the wrong key, the fix is a documentation gap, not a bigger model.
What remains deterministic
- Table naming, layering and certification rules
- Quality rules and the correction register
- Join keys, exclusions and metric formulas in the dictionary
- Row-volume guards and the human gate on every change command
How risk is controlled
- Verify before you write: profiling evidence is required before a column or measure is used
- Change commands are generated, verified and shown to a human before execution
- Correction register designed to survive refresh, with the postmortem recorded
- Row-volume guards in the query builder so an analyst cannot accidentally pull the whole warehouse
- Health review and remediation of finance semantic models
Verification and controls
- Published table-naming reference and a pipeline architecture review against the standard
- Workspace monitoring page for refresh and freshness
- Semantic models developed and health-checked through a modelling toolchain driven from Claude Code
- Reusable skills for hierarchy drift diagnosis and source-truth discovery
Tech stack
Data
- ERP and data-warehouse sources source systems with inconsistent names and rules
- Microsoft Fabric and OneLake the Bronze, Silver and Gold layout with naming and certification rules
- Dataflows and pipelines Bronze landing, then Silver conformance and quality rules
- PySpark, T-SQL, M transformations written after each column is profiled
Application
- Power BI, DAX and Direct Lake semantic models built to consume certified Gold
- Tabular modelling tools semantic-model health checks and fixes
AI at runtime
- Text-to-query over the dictionary AI-generated SQL and DAX read the dictionary, so the model joins on the right key
- Semantic-model change agent drafts the change commands a person approves
AI-assisted development (not at runtime)
- Claude Code drives the modelling toolchain that develops and health-checks semantic models
- Reusable skills shared skill definitions
Adoption and outcomes
A governed data foundation supports consistent definitions, quality review and reusable information. Confidential operational details are withheld. structural
-
Outcome
Bronze to Gold
medallion architecture with certified Gold datasets
realised
-
Engineering
Shared definitions
a governed dictionary supports consistent interpretation
structural
-
Control
Human-gated
every agent-generated change command is verified before it runs
structural
Status In place
What I learned
Encode business rules in a dictionary the model reads, not in prompts someone remembers. When the AI joins on the wrong key, the fix is a documentation gap, not a bigger model.
Related work
Keyboard: [ previous, ] next.