GEHNA.SDM
[ Finance ]

AI-Assisted Bookkeeping and the Judgement That Doesn't Automate

AI is highly accurate at the mechanical layer of bookkeeping. That accuracy can quietly erode the judgement layer it was never designed to replace.

By Gehna Stavonin-de Montagnac8 August 20266 min read

AI can now categorise a transaction, extract data from a receipt, and propose a reconciliation with real accuracy. What it can't do is reliably tell you when something looks wrong in a way that matters — and that distinction, between mechanical accuracy and genuine judgement, is the actual dividing line worth understanding, rather than a vague sense that "some things still need a human."

Fact: AI-assisted bookkeeping tools are pattern-matching systems — they categorise a new transaction based on how similar transactions were treated before, and flag anomalies based on statistical deviation from a business's normal patterns. Both of these are genuinely useful and generally accurate. Neither of them is the same thing as understanding why a transaction happened, or knowing enough about a specific business's actual circumstances to judge whether an unusual pattern is a problem or simply a normal, explainable variation.

What judgement actually means here, concretely

Judgement in bookkeeping and accounting isn't a vague, abstract quality — it's a specific set of habits: noticing when a number that's individually plausible doesn't fit the surrounding context, asking why something changed rather than just recording that it did, and knowing which anomalies are worth a conversation with the client and which are routine. AI can surface the anomaly. It can't reliably tell you which category it falls into, because that requires context — knowledge of the business, its history, its specific circumstances — that isn't fully captured in the transaction data itself.

Analysis: this creates a specific, practical risk worth naming directly: AI's genuine accuracy on the mechanical layer can create false confidence about the judgement layer, precisely because the mechanical output looks so reliable that it's tempting to extend the same trust to the parts that actually require a person thinking it through. A categorisation that's correct the vast majority of the time doesn't mean the remainder doesn't matter — it means that remainder is exactly where judgement is doing real work, and it's the hardest kind of error to catch precisely because it's rare enough that reviewers can get out of the habit of looking closely.

Opinion: the accountants and bookkeepers who get the most value from AI tools long-term aren't the ones who trust the output most, they're the ones who've gotten precise about exactly which parts of their judgement AI can't replace, and who protect deliberate time and attention for exactly those parts rather than letting review quality quietly erode because the tool is right often enough to breed complacency.

Prediction, held loosely: as AI's mechanical accuracy keeps improving, the professional value of judgement doesn't shrink — it concentrates. The routine judgement calls that used to make up a lot of a bookkeeper's day get automated alongside the data entry; what's left is disproportionately the genuinely hard calls, the ones that were always going to need a person. That's a smaller number of decisions, but a higher-stakes one, which changes what "attention to detail" actually means in practice — less about catching typos, more about catching the plausible-looking thing that's actually wrong.

Written by

Gehna Stavonin-de Montagnac

Writing on artificial intelligence, software, automation, business and finance.