In verification processes, time is spent reading heterogeneous documents, not on the final decision.
Anyone who has worked in a compliance function recognizes this. Assessing a case, once all the information is in front of you, takes minutes. Getting all the information in front of you takes hours.
From this follows where it makes sense to intervene, and where intervening would be premature before it’s even inefficient.
Mapping the process and the friction points
A customer due diligence review follows, with variations, the same sequence.
Gathering documentation: company records, articles of association, identity documents, financial statements, declarations, powers of attorney. They come from different sources, in different formats, with different quality.
Extracting the relevant information: legal name and form, ownership structure, powers of representation, beneficial owners, operating address, business activity.
Reconciling across sources: checking that what appears in the company record matches what the customer declares and what shows up in the databases consulted.
Screening against lists and assessing adverse media.
Risk assessment and decision, with documented reasoning.
The friction points are concentrated in the first three stages. Reconstructing an ownership chain spread across multiple corporate layers, in particular, involves sequentially reading multiple documents and manually transcribing data that then has to be cross-checked.
Stages four and five are already largely supported by dedicated tools, and the fifth remains a professional judgment regardless.
Structured extraction and cross-source reconciliation
This is where the most solid contribution lies.
A document intelligence system reads documents that don’t follow a common standard and extracts the same information from them. An Italian company register extract, a foreign registry certificate and a set of articles of association contain the same relevant entities expressed differently: the system maps them into a uniform structure.
From this come three operations that are done by hand today.
Reconstructing the ownership chain. Starting from the corporate documents, the system builds the ownership structure and calculates indirect stakes, flagging the points where the documentation isn’t sufficient to continue.
Cross-checking sources. The system highlights discrepancies: an address that differs between two documents, an expired appointment, a declared fact that finds no confirmation. It doesn’t resolve them. It brings them to the analyst’s attention with the references alongside.
Checking completeness against the customer type and risk level, with a precise list of what’s missing.
Time goes down because the analyst receives a file that’s already been read and structured, with anomalies highlighted, and applies their expertise to the assessment instead of to the gathering.
Traceability is a requirement, not a nice-to-have
It’s worth being blunt on this point, because it’s the criterion that separates a system usable in a control function from one that isn’t.
Every extracted piece of information has to reference the source document and its location within it. Not as an added feature: as the condition that makes the extraction usable at all.
There are three reasons.
The review has to be reconstructible. If, during an internal or regulatory inspection, someone asks on what basis beneficial ownership was determined, the answer has to be a document and a precise location, not a system output.
Human review has to be fast. An analyst who needs to check twenty extracted data points opens twenty references and confirms them. Without a link to the source, they’d have to reread the entire file, and the benefit disappears.
Errors have to be diagnosable. With the source alongside, you can immediately tell a system reading error apart from information that’s genuinely ambiguous in the original document. These are two different problems and require different responses.
On top of this, there’s the log of queries and versions: which body of regulation was in force, which system configuration was active, who consulted what. In a supervised context, this is part of the process documentation.
The effects to measure
Indicators need to be agreed beforehand and recorded against the starting point.
Average review time by risk tier. Broken down by tier, because the effect on ordinary cases differs from complex ones, where it’s typically larger.
Share of files complete on first pass, meaning no need to go back to the customer for missing documentation that wasn’t caught at the start.
Number of discrepancies caught during review versus those caught downstream. An increase in the former, alongside a decrease in the latter, is an improvement even though the first number is rising.
Document search time for the compliance function on internal regulations and policy, when the project also includes semantic search over the regulatory corpus.
A warning: the automation rate isn’t a results indicator. In a control process, a high rate of human intervention on doubtful cases is a sign of correct calibration, not of failure.
What shouldn’t be automated
Risk assessment and the decision remain professional and remain human. This isn’t a matter of style: responsibility for the function can’t be delegated to a system, and an automatic output that proposed a risk classification to confirm would create an anchoring effect that’s hard to correct.
Interpreting adverse media requires context and judgment. A system can gather and summarize; relevance is a judgment call.
Handling suspicious activity reports follows its own procedures with designated responsibilities, which don’t change.
The AI Act framework also has to be considered. Some uses in the financial sector fall under the high-risk categories of Annex III, with obligations that, after the amendments introduced in July 2026, apply from 2 December 2027. A document-support system that doesn’t produce risk assessments sits differently, but the classification has to be made case by case, at the start of the project and with the compliance function.
In closing
Generative AI’s most solid contribution to these processes isn’t in the decision. It’s in preparing for the decision: reading, extracting, reconciling, flagging anomalies.
It’s a less visible contribution than others, and one with a return that’s easier to demonstrate, because the time spent on document-handling stages is measurable and, in many organizations, already tracked today.
If you have a specific process in mind — customer due diligence, credit review, document handling for non-performing positions — we’re available for a conversation about that case: what documents are involved, what traceability requirements apply, and what indicators would be measurable.




