05 · AI operating model · Governed workflows with verifier gates

AI Operations Platform

The operating layer for scaling AI workflows safely.

At a glance

Status
In daily use
What this label is based on

“a one-page view refreshed daily”

The business story

Operational support depended on consolidating status information and making review responsibilities explicit.

AI workflows need clear ownership, evaluation and practical operating controls so people can use and support them.

My contribution

  • Designed and built the operations console end to end: governance doctrine, provenance-per-fact state builder, live overlay rules, artefact guard, staleness model, installer and scheduled rebuild
  • Designed the change-request lane and drift reconciliation for a fleet whose runtime is owned by the business; wrote the rulings log for the owner
  • Own several of the semantic models the agents read, and the upstream pipelines, refresh timing and naming governance
  • Built the evaluation harness (rules first, small-model judge on ambiguous cases) test-first, and set the model-tier and cost rules

The process

  1. Mapping

    Agent manifests and the metrics store mapped as the console's only inputs; the change-request lane and rulings log written for the fleet owner.

  2. Build

    Each fact given a provenance pointer, staleness made schedule-aware, and a build-time guard added that refuses to write a failing artefact.

  3. Validation

    Offline tests including an artefact-invariant suite; independent review strengthened governance and evaluation.

  4. Rollout

    A scheduled daily rebuild promotes the new console only when every live source read succeeded; it is emailed as one file.

How it works

The operating layer around the fleet. A read-only operations console that reads agent manifests and the metrics store as data, attaches a provenance pointer to every derived fact, and is baked into a single self-contained file that can be emailed, plus an optional loopback server that adds a live overlay with a strict rule that a failed read is never shown as a zero. Around it: the governance lane (branch-and-patch change requests, rulings log, drift reconciliation), an evaluation harness for a text-to-query agent, and release and operational gates reused across projects.

An illustrative fleet-health view: each workflow shows schedule-aware freshness and its kill-switch state, and a source that cannot be read shows as unknown, not as healthy. Illustrative states, not live telemetry or a product screenshot.
Illustrative states for the AI Operations Platform, not a product screenshot, live telemetry or a measured result.
01PROBLEM02RULES + DATA03ENGINE04AI05GATE06WORKFLOW07OUTCOMEAI workflowsneeding sharedoversightAgent contract,principles,rulingsDeterministicfigures,sentence-IDcommentaryModels order andphrase, nevercomputeKill switch,release tag,drift checkConsole: health,staleness, whereto change itFleet triage inone page, daily 01PROBLEM02RULES + DATA03ENGINE04AI05GATE06WORKFLOW07OUTCOMEAI workflows needing sharedoversightAgent contract, principles,rulingsDeterministic figures,sentence-ID commentaryModels order and phrase, nevercomputeKill switch, release tag, driftcheckConsole: health, staleness,where to change itFleet triage in one page, daily
  • Deterministic stage
  • AI stage
  • Gate: human decision point
AI role Models write commentary. They never decide a figure.
  1. 01 Problem

    AI workflows needing shared oversight

    A hand-kept status table had already drifted from reality.

  2. 02 Rules + data

    Agent contract, principles, rulings

    In the fleet, owned by its runtime owner: a written agent contract and principles document, with rulings and change requests recorded. I designed the change-request lane and the rulings log.

  3. 03 Engine

    Deterministic figures, sentence-ID commentary

    In the fleet: every business figure is computed in plain code or published measures; the model receives pre-written sentences and returns their IDs.

  4. 04 AI

    Models order and phrase, never compute

    In the fleet: a single shared model call site logs prompt hashes; an unknown sentence ID is dropped and the report is complete with no model response at all.

  5. 05 Gate

    Kill switch, release tag, drift check

    In the fleet: a kill switch per agent, tagged-release deploys and a daily drift check. The console classifies that switch state as on, off, unknown or incomplete rather than guessing.

  6. 06 Workflow

    Console: health, staleness, where to change it

    A read-only console built from manifests and the metrics store, with a provenance pointer on every fact, emailed as one file.

  7. 07 Outcome

    Fleet triage in one page, daily inferred

    Fleet triage in one page, rebuilt daily; the time saved is inferred, not measured.

Step by step
  1. A build step reads every agent manifest and the metrics store as JSON and read-only SQL, never by importing agent code.
  2. Each fact carries a provenance pointer. Unknown is null. Staleness is schedule-aware, so a weekday agent is not stale on Saturday.
  3. Kill-switch state is classified as on, off, unknown or incomplete rather than guessed.
  4. A build-time guard with a denylist of write tokens and an allowlist of permitted read paths refuses to write the artefact if it fails.
  5. A scheduled daily rebuild promotes the new console only when every live source read succeeded; otherwise the previous artefact is restored.
  6. Change requests travel to the runtime owner as patch bundles with a sign-off checklist; the packager refuses sensitive paths and scans for secret-shaped strings.
  7. In the fleet itself (owned by its runtime owner, surfaced by the console): deterministic figures, commentary by sentence ID from one logged call site, tagged-release-only deploys, a daily drift check and a per-agent kill switch.

AI and engineering judgement

Where AI is used

In the fleet (owned by its runtime owner), models order and phrase pre-written sentences; they never compute a figure, because every figure is computed in plain code or published measures and the model returns only the IDs of the sentences to use. In the console I built, an advisory guide panel answers questions about the fleet and is never a source of figures. In the evaluation harness, a small model judges only the ambiguous cases a rule-based rubric cannot settle, with a hand-labelled calibration set and a cost boundary asserted in tests.

What remains deterministic

  • Every business figure in every report
  • Staleness, kill-switch and drift classification
  • The build-time read-only guard and the daily promotion gate
  • The rule-based rubric that runs before any model judge

How risk is controlled

  • Decide what an agent may never do before what it may do: a leadership assistant maps questions to a fixed set of certified figures and cannot write a query
  • In the fleet (surfaced by the console, not owned by it): a kill switch per agent that trusts only explicit values, with lifecycle kept as a separate axis
  • In the fleet: production deploys only from an annotated tag on a green main branch and a published release asset, dispatched by a person who types the tag twice
  • In the fleet: continuous integration that fails when a waiver’s condition clears, so stale exceptions cannot survive
  • The console reads no secrets at build or run time; a model key, if entered, lives in page memory for the session only
  • Release-readiness, schema-migration, secrets-lifecycle and operational-readiness gates that halt for human sign-off
Verification and controls
  • Offline tests on the console, including an artefact-invariant suite and schedule-parsing and live-overlay suites with cloud calls stubbed
  • Cross-model review: an independent second model family reviews diffs before they ship; a council pattern with blind cross-ranking pressure-tests decisions
  • Independent review strengthened governance and evaluation
  • A portable operating-system kit (manual, roles, playbooks, skill registry, memory seed) that stands the setup up in a fresh account

Tech stack

Data

  • Microsoft Fabric and Power BI semantic models the semantic models the agents read, several of which Tommy owns along with their upstream pipelines

Application

  • Python (stdlib server and build) builds the read-only console and its optional loopback live overlay

AI at runtime

  • Advisory guide panel answers questions about the fleet; every figure shown comes from the console build, not the panel
  • Small-model judge judges only the ambiguous cases a rule-based rubric cannot settle
  • Fleet commentary models in the fleet, owned by its runtime owner: return the IDs of pre-written sentences; the figures are computed in code

Automation

  • Windows Task Scheduler and PowerShell the scheduled daily rebuild with its quality gate

Deployment

  • Azure Functions and App Configuration fleet runtime platform, owned by its runtime owner; the console reads it as data
  • Application Insights runtime telemetry, read by the console as data rather than re-measured

Testing and CI

  • pytest the console's test suite

AI-assisted development (not at runtime)

  • Claude Code, skills and MCP the build environment, with release and operational gates reused across projects
  • Second-model-family review reviews diffs before release
  • Council review a structured second opinion on decisions

Adoption and outcomes

Workflow oversight is supported by a one-page view refreshed daily and documented review controls. Fleet scale and commercial impact are not disclosed. inferred

  • Outcome

    Shared oversight

    AI workflows are visible through a governed operating view

    current

  • Engineering

    Provenance per fact

    unknown is null; a failed read is never shown as a zero

    structural

  • Control

    Read-only artefact

    the emailed console is built read-only and verified at build time

    structural

Status In daily use

What I learned

Decide what an agent is never allowed to do before deciding what it can do. The leadership assistant is safe because it cannot write queries; the fleet is safe because every agent has an off switch someone owns.

Related work