01 · Finance operations · Chain with a confidence gate
AP Intelligence
Human-reviewed document intelligence for finance operations.
At a glance
- Status
- In production
What this label is based on
“pilot to production-hardened build”
The business story
The rules for checking each supplier’s documents were remembered rather than written down.
Instead of transcribing, the team now reviews. Each document is classified, extracted and checked against written rules, and only the uncertain rows reach a person before anything is exported.
My contribution
- Designed the end-to-end pipeline and the validation rulebook with the finance team
- Set the governance: action-cost gates, data-residency options, release standard
- Directed AI coding agents under that governance and reviewed their output
- Ran user acceptance testing with the accounts-payable team and handed over with process maps and training
The process
-
Mapping
Each supplier format described once, and the validation rules (totals, running balances, GST, dates, required fields, sign rules) written down, not remembered.
-
Co-design
The pipeline stages, from intake to finalise, were agreed with the finance team.
-
Build
Iterative construction separated extraction, validation and review responsibilities.
-
Validation
Golden-fixture evaluation against hand-labelled documents, plus continuous integration with a deploy-parity smoke job and a nightly UI job.
-
Rollout
Acceptance testing with amber and red rows reviewed in the queue before any export.
-
Handover
Process maps and training, with a batch runner for unattended processing.
How it works
A pipeline that turns supplier documents into validated, reviewable structured data for reconciliation. It ships as a self-contained tool for non-technical users, who install nothing, with a batch runner for unattended processing and an evaluation command for new document formats.
- Deterministic stage
- AI stage
- Gate: human decision point
-
01 Problem
Supplier PDFs and statements
Supplier statements re-keyed by hand every month, in several inconsistent layouts.
-
02 Rules + data
Vendor formats and validation rulebook
Each supplier format is described once. Validation rules (totals, running balances, GST, dates, required fields, sign rules) are written down, not remembered.
-
03 Engine
Classify, extract, post-process
A chain: detect digital, scanned or spreadsheet; classify the format; extract rows; apply business rules and overrides.
-
04 AI
Model extraction at temperature zero
The extractor is either a rule engine or a model behind one provider seam, so the model can be swapped or run locally for data residency.
-
05 Gate
Deterministic validation and confidence gate
Every row is re-checked deterministically and banded green, amber or red. Amber and red go to a person.
-
06 Workflow
Reviewer queue, then explicit finalise
Reviewers edit in a queue; edits are re-validated and audited; a run must be finalised before export.
-
07 Outcome
Reviewed structured output realised
Validated output supports review, with an audit trail on edits.
Step by step
- Intake: documents arrive by email or file drop and are detected as digital, scanned or spreadsheet.
- Classification: each document is matched to a known supplier format, or routed for template bootstrapping.
- Extraction: a rule-based engine or a model at temperature zero extracts rows to typed JSON behind one provider seam.
- Post-processing: business rules, supplier-name overrides and payment-row filtering.
- Validation: balance and running-total cross-checks, required fields, sign rules, date and number checks.
- Confidence gate: every row is banded green, amber or red. Flagged rows go to a review queue.
- Review and finalise: edits are re-validated and audited, and a run must be finalised before the clean file is produced.
AI and engineering judgement
Where AI is used
Extraction of rows from messy layouts, where a model copes with variation that rules alone do not. A vendor template bootstrapper uses a handful of samples from a new supplier to generate a reusable extraction template. Nothing else in the pipeline depends on a model.
What remains deterministic
- Document and format classification rules
- All validation: totals, running balances, GST, required fields, sign and date checks
- Confidence banding thresholds
- The review, edit, re-validate and finalise workflow
- Export formats and the audit log
How risk is controlled
- Stub mode by default; live mode refuses to run without required configuration
- Batch cap, single-writer run lock and an idempotent content-hash manifest so re-runs cannot duplicate work
- Quarantine of oversized or unexpected files
- Written rule: a person reviews anything supplier-facing or payment-related before it leaves
- Append-only run log written by both the app and the headless path
- Secrets kept out of the repository; the example configuration holds placeholders only
Verification and controls
- Golden-fixture evaluation: field match, numeric error and row recall against hand-labelled documents, so accuracy is a published number
- Continuous integration: lint, format check and tests across a language-version matrix, a deploy-parity smoke job and a nightly UI job
- Single provider seam: multiple model providers and a local open-weights option for on-premises residency
- Decision log, learnings and open-issues files kept in the repository so work survives across sessions
Tech stack
Data
- pandas / openpyxl spreadsheet intake and the clean structured export
- PDF extraction reads supplier PDFs ahead of classification
Application
- Python one pipeline from intake to finalise, plus a batch runner for unattended processing
- Streamlit the self-contained interface for non-technical users, who install nothing
AI at runtime
- Multi-provider model router one provider seam, so the model can be swapped or run locally for data residency
Testing and CI
- pytest and golden fixtures the test suite and the accuracy evaluation
- GitHub Actions runs the automated checks
AI-assisted development (not at runtime)
- AI coding agents directed under the project governance, with their output reviewed
Adoption and outcomes
The workflow supports validated document handling and a traceable review process. Commercial impact is not disclosed. realised
-
Outcome
Reviewed outputs
less repetitive document handling, with validation and an audit trail
realised
-
Engineering
Validated build
pilot to production-hardened build
realised
-
Control
Green / amber / red
confidence gate on every extracted row
structural
Status In production
What I learned
Separate extraction from validation, and validation from approval. Each layer has its own test set, so a new document format is caught by that layer's tests instead of degrading the rest of the pipeline.
Related work
Keyboard: [ previous, ] next.