Human in the Loop, and Where to Actually Check AI Bookkeeping Work
Reviewing everything is not a control, it is the absence of one: uniform review becomes rubber-stamping within a fortnight. Here is what an Australian practice should always check, what it should sample, and what it should never hand over at all.
Short Answer
Review 100% of anything that leaves the firm, anything lodged, anything above materiality, anything the system flagged, and every output in a new client's first cycle. Sample routine categorisation and reconciliation at a rate set by measured accuracy. Never automate judgement, advice, or any output that cannot cite a source record. Uniform review is not stricter, it is a control that fails quietly.
Last reviewed: September 2026
Key takeaways
- "Human in the loop" has no operational meaning until a firm writes down which outputs it covers, and ASIC has documented four incompatible readings of the phrase in the Australian market.
- Checking every AI output deletes the time saved and produces rubber-stamping, degrading into no review at all within a fortnight.
- TPB(GS) 55/2026 does not require uniform review: it says there is no set formula for reasonable care.
- Auditing Standard ASA 530 already answers "how much should I check": it defines sampling as testing less than 100% of a population.
- A review design needs three tiers, always review, sample review and never automate, with a written threshold and rate per output type, and the half vendors skip is measuring the reviewer so the rate moves on evidence rather than comfort.
Every vendor pitch and every regulator page now carries the phrase "human in the loop", and almost none say which loop, which human, or how often. That leaves Australian bookkeepers between two bad options: check everything, which deletes the reason for deploying dedicated AI at all, or check nothing, which is the risk the phrase was invented to manage. The answer is neither, and it is a design rather than a slogan.
What does "human in the loop" actually mean in Australian accounting?
Four different things, according to the only Australian regulator to have counted. In REP 798 Beware the gap: Governance arrangements in the face of AI innovation, published 29 October 2024, ASIC reviewed 624 AI use cases across 23 AFS and credit licensees as at December 2023, asking each whether its models ran with a human in the loop. Most said yes. ASIC then recorded what "yes" meant: "this ranged from using the model's output to inform a human decision-maker, to referring exceptions to a human for review, to having human involvement in training and to periodically testing the model's operation."
Those are not variations on a theme. Referring exceptions to a human and periodically testing a model are different controls with different failure modes, so a firm that says "we keep a human in the loop" without saying which has described a feeling, not a procedure. ASIC's question back to the market is the one to steal: are you clear on what human oversight you expect, and do you have procedures for when things go wrong?
ASIC also names the trap: some licensees identified the risk of an incorrect model output but treated it as handled because "the consumer could contest it, or staff could override it". Naming a human as the backstop is not the same as showing the human catches anything.
Why is "review everything" not a control?
Because attention is finite and uniform review spends it on the wrong items. A bookkeeper handed 900 AI-coded transactions does not review 900; they review the first forty, notice the system is getting it right, and start clicking. Within a fortnight the approval step is a keystroke: a control on the org chart and none in the ledger, worse than none at all, because the false assurance stops anyone designing a real one.
The second cost lands sooner. Uniform review consumes exactly the hours the deployment was meant to release, so the business case never arrives and the tool gets blamed for it, the failure pattern set out in the post on why AI saves accountants no time.
The fear driving over-review is real. ICAS-commissioned research by Alliance Manchester Business School and Aston Business School, published 17 March 2026 from a survey of more than 200 accounting professionals at a UK mid-tier firm, found 72% fear generative AI could produce errors or reach incorrect decisions. That is UK sentiment data, not an Australian error rate, but a fear answered with a slogan stays a fear.
Does the Tax Practitioners Board require you to review every AI output?
No, and it says so almost in those words. TPB(GS) 55/2026 The use of Artificial Intelligence and the Code of Professional Conduct, issued 22 July 2026, requires that practitioners "should verify and review AI generated content for accuracy throughout each step of the workflow", adding that "Each of these steps should be documented." On how much checking constitutes reasonable care it is explicit: "there is no set formula", and the answer "will depend on an examination of all the circumstances".
The circumstances it lists are a review design in miniature: the nature and scope of the service; whether the practitioner checks the output before relying on it; and whether AI has been used for tasks which should be independently verified, particularly outside their expertise. That is a risk-weighted instruction. Disclosure under the same guidance is covered in the breakdown of what tax and BAS agents must disclose.
Government guidance points the same way. The National AI Centre's Guidance for AI adoption: implementation guidance, current PDF published 5 May 2026, devotes its sixth essential practice to "Maintain human control": accountability, the ability to intervene, training so the overseer understands each system's limitations and failure modes, and alternative pathways if it goes offline. Guardrail 5 of the earlier Voluntary AI Safety Standard reads the same. Neither says review every output. Both say know who is accountable, know when to step in, and be able to.
Is sampling AI work instead of checking all of it defensible?
Yes, and the profession settled this long before AI arrived. Auditing Standard ASA 530 Audit Sampling, approved by the AUASB on 3 March 2020 and operative for financial reporting periods beginning on or after 15 December 2021, defines audit sampling as "the application of audit procedures to less than 100% of items within a population of audit relevance such that all sampling units have a chance of selection in order to provide the auditor with a reasonable basis on which to draw conclusions about the entire population".
An Australian auditor signing an opinion does not examine every transaction, and nobody calls that negligent. The rigour is in the discipline around the sample, and ASA 530 sets out four requirements worth lifting into an AI review policy: the sample shall reduce sampling risk to an acceptably low level; every item shall have a chance of selection; the auditor shall investigate the cause of any deviation found; and misstatements found shall be projected to the population.
That last one separates sampling from spot-checking, and almost every AI deployment omits it. If a bookkeeper reviews 50 of 900 AI-coded transactions and finds two miscodings, the firm has not found two errors. It has an estimate of roughly 36 across the population, and a decision about whether that is tolerable.
Which outputs get reviewed every time, which get sampled, and which never get automated?
Three tiers, written per output type with a threshold and a rate. An output that cannot be placed in one of them has not been thought about yet.
Tier 1, always review, no exceptions
Anything leaving the firm, anything lodged, anything above the client's materiality threshold, anything the system flagged as uncertain, and every output in a new client's first cycle. The last two carry the most weight: an ignored flag is worse than no flag, and accuracy on one chart of accounts predicts nothing about the next one.
Tier 2, sample review
Routine, high-volume, low-consequence AI bookkeeping work: categorisation, reconciliation matching, supplier statement matching, standard journals. Each gets a written rate that starts at 100%, falls while measured accuracy holds, and rises the moment it does not. That rate is a number in a policy document, not a habit in someone's head.
Tier 3, never automate
Judgement calls, advice, and anything the system cannot tie to a source record. The third is the one firms miss: an output that cannot point at the invoice, bank line or clause it came from cannot be checked, only believed, which produces the failures in the piece on hallucinations in financial reports and sits within the limitations of AI employees.
| Output type | Review mode | Trigger or threshold | Indicative steady-state rate |
|---|---|---|---|
| BAS, IAS or any lodgment | Always review | Every lodgment, no threshold | 100% |
| Client-facing email, report or statement | Always review | Anything leaving the firm | 100% |
| Any transaction above materiality | Always review | Dollar threshold set per client | 100% |
| Output the system flagged as uncertain | Always review | Any escalation raised by the system | 100% |
| Anything in a new client's first cycle | Always review | First full period, per workflow | 100% |
| Routine transaction categorisation | Sample review | Steps down only while measured error stays in tolerance | Set by measurement, reset on any change |
| Bank reconciliation matching | Sample review, plus 100% of unmatched items | Every exception reviewed regardless of rate | Set by measurement |
| Deductibility and treatment judgements | Never automate | Facts not in the ledger | Not applicable, human decides |
| Advice to a client | Never automate | Any output shaping a decision | Not applicable, human decides |
| Any output with no citable source record | Never automate | System cannot name its source record | Not applicable, blocked by design |
The rates left as "set by measurement" are deliberate: a universal number would repeat the vendors' mistake, since the right rate for a tidy single-entity client is not the right rate for a multi-entity group mid-restructure.
How do you set the sample rate without guessing?
By measuring the reviewer, the half of human-in-the-loop that almost no vendor documentation mentions. The output of a review is not an approval, it is a data point: items examined, errors found, what kind, what caused them. Recorded per sample and per reviewer, that turns the rate into an evidence-based setting instead of a vibe.
Three numbers matter from week one. The detected error rate per sample, which sets the rate. The cause mix, because a miscoding from an ambiguous supplier name needs a rule change while a fabricated figure needs a halt. And the reviewer's detected rate against an independent re-check of the same population, the only reliable rubber-stamping detector: when detected errors fall to zero while a second pass still finds them, shorten the queue rather than exhort the reviewer.
All of this assumes the actions were logged well enough to re-check, and the test for that is in the article on proving what your AI actually did. A review of an unlogged action is an opinion about a memory.
What should the first 90 days of review look like?
- Days 1 to 5: write the tier table before switching anything on. Assign every output type to always review, sample review or never automate, and set the materiality threshold per client. An output nobody can classify does not go live.
- Days 1 to 5: define what counts as an error. Separate a wrong account code, a wrong GST treatment, a wrong amount and a fabricated figure. Without this the error rate means nothing later.
- Days 6 to 30: run one workflow, one client, at 100% review. Not the whole practice, not the messiest client. Record every error, its type and its cause.
- Day 30: hold the first measurement review. Produce the error rate, the cause mix and the hours actually spent reviewing. Step the rate down only if the rate is inside tolerance and the causes are understood.
- Days 31 to 60: add a second client and workflow, each starting at 100%. The rate earned on client A does not transfer to client B, and this is where firms generalise a result they never measured.
- Day 60: introduce the independent re-check. A second person re-reviews a slice of a sample another reviewer already approved. This measures the review, not the AI.
- Days 61 to 90: make the reset rules automatic. Any error above materiality, any software or model version change, any material change in the client's business, and the rate returns to 100%.
- Day 90: publish the numbers into the AI policy. Tier table, current rates, measured error rates and reset rules become a document a reviewer or the TPB can read, structured as in the AI policy template for Australian accounting firms.
Ninety days is a floor: a monthly client needs one complete month-end at full review per workflow before any rate moves, and quarterly BAS work needs a quarter. Staffing and sequencing are in the AI employee onboarding checklist.
How does Agentive configure review for Australian practices?
With escalation rules rather than a blanket approval queue. An approval queue asks a human to confirm everything and gets confirmed by reflex; an escalation rule asks the system to stop on a defined condition, so the items reaching a person are the items that needed one. In an Agentive deployment those conditions are written per client: a dollar threshold, a GST treatment not seen before on this ledger, a supplier with no coding history, a variance against the prior period, an output with no source record to cite. Everything else proceeds into the sample pool.
The observation from configuring this for Australian finance teams is that firms relax too early, and for the wrong reason. The trigger should be a measured error rate on that client's own data across a complete cycle. What it usually is instead is confidence: the team watches the system be right for a fortnight and stops looking. A rate never derived from a number cannot be defended to a reviewer, a client or the Tax Practitioners Board. The corollary is worth saying plainly: an Agentive deployment is more work than the status quo in its first month, and any vendor promising otherwise is describing month six.
Agentive runs single-tenant on AWS Sydney, all inference inside Australia, client data never leaving Australian borders and never used to train a model, aligned to APRA CPS 234, ASIC RG 255 and TPB obligations including TPB(GS) 55/2026. Every action is logged as it happens, with the source record read and the escalation state, which is what makes the independent re-check possible at all. Where escalations sit relative to the client relationship is taken up in the comparison of client-side and back-office agents.
What this design does not fix
A sampled review still misses things, by construction. ASA 530 names it: sampling risk, the risk that a conclusion drawn from a sample differs from the one testing the whole population would produce. Risk-weighted review swaps an unknown exposure for a known, bounded, measurable one, and that trade should be documented rather than discovered later.
Nor does review transfer accountability. TPB(GS) 55/2026 states that tax practitioners remain ultimately responsible for the services they provide, so no sample rate changes who signs. That is why, for bookkeepers running this at volume, the always-review tier is written in absolutes rather than percentages.
Write the rate down, then move it on evidence
This cluster opened on a regulator's disclosure obligations and arrives ten posts later at the same place: the firms making AI work in Australian practice turned a phrase into a written procedure with numbers in it. "Human in the loop" is not a control. A tier table, a materiality threshold, a sample rate, a measured error rate and a reset rule are a control, and they fit on one page.
So pick one client and one workflow this week, classify every output into the three tiers, and review at 100% for a cycle while counting. At the end you have a number, and the number sets the rate. That is the difference between an AI Operation Engine a practice can defend and a tool it merely hopes about.
General information for Australian practices, not professional, legal or tax advice. Regulatory positions are stated as at 19 September 2026 and should be verified against the primary sources linked above. Sample rates and thresholds are decisions for each firm and are illustrative here.
Bring Us Your Ledger and We Will Design the Review Around It
Agentive builds dedicated AI for Australian bookkeepers, accountants and BAS agents, single-tenant on AWS Sydney, all inference inside Australia, client data never used to train a model, and a per-action log written at the time of the action. Bring one client file to the call and we will map which outputs get reviewed every time, which get sampled, and which never leave a human.