Agentic tools are showing up in recon queues, forecast refreshes, and exception triage. Fine. The question that matters on an Australian finance desk isn't "can it run?" It's "can we reconstruct the decision when someone asks?"

Audit, internal review, a bank, a board member — same test. What did the agent see? What rule or threshold did it apply? What did it propose? Who approved? Where is that logged?

I'm pro-foundation and pro-sequencing. Useful automation with accountable humans. Soft on vendors: the product can be excellent and still fail your control file if the trail isn't there.

Why this sits at the top of the pile

Workday's ANZ work on realising ROI from AI agents in finance puts the operating reality in plain numbers: 89% say data isn't ready for agents; 87% plan to retain human-in-the-loop; 77% aren't seeing ROI; 95% put risk and compliance among their top AI priorities.

That pattern matches what controllers and recon leads already know. If the stack is fragmented, an agent mostly speeds up the same gaps. If risk and compliance are the priority, the path is clear: design for reconstructability first, then widen scope where outcomes are measurable.

Where it fits naturally: Deloitte's 2Q 2026 CFO Signals (23 July 2026) had 43% of CFOs confident in their AI governance, 59% naming deploy-fast-versus-risk as a top challenge, and only 19% saying the CFO owns AI governance. Governance lag isn't a footnote — it's the gap this checklist is meant to close.

Reconstruct-the-decision: minimum trail

Treat this as a production gate, not a slide.

  • Inputs named: source systems, cut-off times, master data version, and which fields the agent was allowed to read
  • Logic visible: rule, threshold, or model version that drove the proposal — in language a FC can explain
  • Proposal vs authority: agent drafts or flags; named human approves cash movement, master changes, or anything relied on for board/bank/audit
  • Exception queue: owner, SLA, and escalation — not a shared inbox that "someone" watches
  • Immutable log: enough to replay the decision six months later without calling the vendor's support line

If any of those five are fuzzy, keep the agent in propose-only mode on low-risk workflows until they're not.

Where agents can propose first

High-volume recon match against known rules. Variance explanation drafts from reconciled actuals. Anomaly flags against a defined threshold. Forecast refresh assembly from owned inputs.

Where a human must approve: payment runs, facility draws, entity/bank mapping changes, AML/KYC-adjacent exceptions, and any output that will be relied on externally.

Desk test before go-live

Pick one workflow. Run it for two weeks with the trail on. Ask your controller: "Would you defend this in a walkthrough?" If the answer is hedged, fix the trail — don't widen the agent's scope.

Same test works for FP&A assistants and recon bots. If the FC can't explain the proposal in two minutes without calling the vendor, it isn't ready for anything that touches the board pack or a bank conversation. Brisbane desks and national teams face the same bar: useful automation, accountable humans, reconstructable calls.

CTA: Building the same guardrails? Compare notes with peers at financesignal.ai — practical control thinking from the desk, not a product pitch.