01 · Finance operations · Chain with a confidence gate

AP Intelligence

Human-reviewed document intelligence for finance operations.

Conceptual illustration of paper documents becoming organised records, with a sheet set aside for review.
Conceptual illustration

At a glance

Status
In production
What this label is based on

“pilot to production-hardened build”

The business story

The rules for checking each supplier’s documents were remembered rather than written down.

Instead of transcribing, the team now reviews. Each document is classified, extracted and checked against written rules, and only the uncertain rows reach a person before anything is exported.

My contribution

  • Designed the end-to-end pipeline and the validation rulebook with the finance team
  • Set the governance: action-cost gates, data-residency options, release standard
  • Directed AI coding agents under that governance and reviewed their output
  • Ran user acceptance testing with the accounts-payable team and handed over with process maps and training

The process

  1. Mapping

    Each supplier format described once, and the validation rules (totals, running balances, GST, dates, required fields, sign rules) written down, not remembered.

  2. Co-design

    The pipeline stages, from intake to finalise, were agreed with the finance team.

  3. Build

    Iterative construction separated extraction, validation and review responsibilities.

  4. Validation

    Golden-fixture evaluation against hand-labelled documents, plus continuous integration with a deploy-parity smoke job and a nightly UI job.

  5. Rollout

    Acceptance testing with amber and red rows reviewed in the queue before any export.

  6. Handover

    Process maps and training, with a batch runner for unattended processing.

How it works

A pipeline that turns supplier documents into validated, reviewable structured data for reconciliation. It ships as a self-contained tool for non-technical users, who install nothing, with a batch runner for unattended processing and an evaluation command for new document formats.

01PROBLEM02RULES + DATA03ENGINE04AI05GATE06WORKFLOW07OUTCOMESupplier PDFsand statementsVendor formatsand validationrulebookClassify,extract,post-processModel extractionat temperaturezeroDeterministicvalidation andconfidence gateReviewer queue,then explicitfinaliseReviewedstructuredoutput 01PROBLEM02RULES + DATA03ENGINE04AI05GATE06WORKFLOW07OUTCOMESupplier PDFs and statementsVendor formats and validationrulebookClassify, extract, post-processModel extraction at temperaturezeroDeterministic validation andconfidence gateReviewer queue, then explicitfinaliseReviewed structured output
  • Deterministic stage
  • AI stage
  • Gate: human decision point
AI role AI extracts. Rules validate. People approve.
  1. 01 Problem

    Supplier PDFs and statements

    Supplier statements re-keyed by hand every month, in several inconsistent layouts.

  2. 02 Rules + data

    Vendor formats and validation rulebook

    Each supplier format is described once. Validation rules (totals, running balances, GST, dates, required fields, sign rules) are written down, not remembered.

  3. 03 Engine

    Classify, extract, post-process

    A chain: detect digital, scanned or spreadsheet; classify the format; extract rows; apply business rules and overrides.

  4. 04 AI

    Model extraction at temperature zero

    The extractor is either a rule engine or a model behind one provider seam, so the model can be swapped or run locally for data residency.

  5. 05 Gate

    Deterministic validation and confidence gate

    Every row is re-checked deterministically and banded green, amber or red. Amber and red go to a person.

  6. 06 Workflow

    Reviewer queue, then explicit finalise

    Reviewers edit in a queue; edits are re-validated and audited; a run must be finalised before export.

  7. 07 Outcome

    Reviewed structured output realised

    Validated output supports review, with an audit trail on edits.

Step by step
  1. Intake: documents arrive by email or file drop and are detected as digital, scanned or spreadsheet.
  2. Classification: each document is matched to a known supplier format, or routed for template bootstrapping.
  3. Extraction: a rule-based engine or a model at temperature zero extracts rows to typed JSON behind one provider seam.
  4. Post-processing: business rules, supplier-name overrides and payment-row filtering.
  5. Validation: balance and running-total cross-checks, required fields, sign rules, date and number checks.
  6. Confidence gate: every row is banded green, amber or red. Flagged rows go to a review queue.
  7. Review and finalise: edits are re-validated and audited, and a run must be finalised before the clean file is produced.

AI and engineering judgement

Where AI is used

Extraction of rows from messy layouts, where a model copes with variation that rules alone do not. A vendor template bootstrapper uses a handful of samples from a new supplier to generate a reusable extraction template. Nothing else in the pipeline depends on a model.

What remains deterministic

  • Document and format classification rules
  • All validation: totals, running balances, GST, required fields, sign and date checks
  • Confidence banding thresholds
  • The review, edit, re-validate and finalise workflow
  • Export formats and the audit log

How risk is controlled

  • Stub mode by default; live mode refuses to run without required configuration
  • Batch cap, single-writer run lock and an idempotent content-hash manifest so re-runs cannot duplicate work
  • Quarantine of oversized or unexpected files
  • Written rule: a person reviews anything supplier-facing or payment-related before it leaves
  • Append-only run log written by both the app and the headless path
  • Secrets kept out of the repository; the example configuration holds placeholders only
Verification and controls
  • Golden-fixture evaluation: field match, numeric error and row recall against hand-labelled documents, so accuracy is a published number
  • Continuous integration: lint, format check and tests across a language-version matrix, a deploy-parity smoke job and a nightly UI job
  • Single provider seam: multiple model providers and a local open-weights option for on-premises residency
  • Decision log, learnings and open-issues files kept in the repository so work survives across sessions

Tech stack

Data

  • pandas / openpyxl spreadsheet intake and the clean structured export
  • PDF extraction reads supplier PDFs ahead of classification

Application

  • Python one pipeline from intake to finalise, plus a batch runner for unattended processing
  • Streamlit the self-contained interface for non-technical users, who install nothing

AI at runtime

  • Multi-provider model router one provider seam, so the model can be swapped or run locally for data residency

Testing and CI

  • pytest and golden fixtures the test suite and the accuracy evaluation
  • GitHub Actions runs the automated checks

AI-assisted development (not at runtime)

  • AI coding agents directed under the project governance, with their output reviewed

Adoption and outcomes

The workflow supports validated document handling and a traceable review process. Commercial impact is not disclosed. realised

  • Outcome

    Reviewed outputs

    less repetitive document handling, with validation and an audit trail

    realised

  • Engineering

    Validated build

    pilot to production-hardened build

    realised

  • Control

    Green / amber / red

    confidence gate on every extracted row

    structural

Status In production

What I learned

Separate extraction from validation, and validation from approval. Each layer has its own test set, so a new document format is caught by that layer's tests instead of degrading the rest of the pipeline.

Related work