Futura AI
it
All articles
  • Data
  • Finance

How we measure 92% automatic payment reconciliation

92% of transactions reconciled automatically, on a sample of roughly 15,000 manually verified payments. The method behind the number, and why raising automation can worsen precision.

by Daniele Grotti3 min read
Futura AI — How we measure 92% automatic payment reconciliation

An incoming payment arrives with a free-text payment reference. The system has to work out which client it belongs to — the NDG, in banking terms — by reading references written differently by different senders.

The number we publish is 92%: the share of transactions reconciled automatically on a comparison sample. It’s worth explaining how it was measured, because this is a case where improving one indicator worsens another, and the number alone doesn’t tell that story.

The sample, not the whole flow

The real scope is on the order of hundreds of thousands of transactions a year. The sample used for measurement is smaller and stricter: roughly 15,000 transactions with the correct NDG already attributed manually by operators, used as the comparison reference. It’s work already done by humans, not an approximation.

The most recurring reference types are covered in full. The less regular structures are being extended progressively: the 92% describes the scope covered today, not the entire payment flow.

Two numbers that need to be read together

Automatic reconciliation isn’t the only quantity measured. What also matters is how many of those automatically attributed payments end up on the correct NDG.

The two move in opposite directions as the confidence threshold changes: raising the automation share worsens attribution precision, lowering it improves precision but leaves more cases to an operator. It’s the same trade-off described for semantic search on another project, applied here to a different problem: not document retrieval, but matching against a client registry.

The two possible errors don’t weigh the same. A payment left to an operator is a processing cost, recoverable. A payment attributed to the wrong position is an error that propagates downstream, and it weighs more. The thresholds are calibrated on this asymmetry: the system prefers routing an uncertain case to a person over risking the wrong attribution.

The 8% left to an operator

One transaction in thirteen, in the measured sample, isn’t reconciled automatically.

For each one, the operator receives the possible candidates already ranked, with the reasoning behind each. The human decision starts from work already done, not from a line of text to interpret from scratch. Every automatic attribution, in the same way, keeps a record of the criterion that produced it and can be reconstructed after the fact — a requirement that, in an audit context, matters as much as accuracy itself.

What isn’t public yet

The rate of incorrect attribution on the verified sample and the measurement’s reference period aren’t published yet: they’re the next figures to confirm with the client, not a detail we overlooked.

In closing

A number like “92% automatic reconciliation” invites you to read it as a pure success rate. It isn’t: it’s a point chosen on a curve that trades automation for precision, calibrated on the asymmetry between the two possible errors. Publishing the curve, not just the point, is what makes the number verifiable instead of promotional.

The full case study describes architecture and thresholds. If you manage a reconciliation process with non-standardized payment references, we’re available to talk through your sample and the thresholds that would make sense for your case.

If this topic touches a real process in your organization, evaluate it with a focused AI Assessment.

Request an AI Assessment