In almost every project we analyze, the bottleneck isn’t the technology. It’s the state of the company’s documentation.
It’s not a flattering conclusion. Meetings are happy to debate which model to adopt, much less happy to discuss how many versions of the same procedure coexist in the archive. Yet it’s the second topic that determines how the project turns out.
What we find when we open an archive
The patterns repeat with surprising regularity, in public bodies and industrial companies alike.
Duplicate documents scattered across multiple folders, with different names and near-identical content. Nothing indicates which copy is the authoritative one.
Outdated versions that remain accessible right next to the current ones. The correct version is known to whoever has worked in that office for years, not to whoever consults the archive.
Scans of uneven quality: historical documents captured at low resolution, hand-filled forms, photographed attachments.
Inconsistent formats for the same type of information. The same technical spec exists as a PDF, a spreadsheet, and a table inside a slide deck.
And a significant share of knowledge that isn’t written down anywhere. The exceptions, the edge cases, the established practices: they live in the experience of a few people.
None of these conditions stop the organization from working. People compensate. That’s exactly why the problem stays invisible until someone tries to automate it.
Why an excellent model on inconsistent data produces inconsistent answers
A system that answers by drawing on company documents has no way of knowing which version is the right one, if nothing in the document itself says so.
If three versions of a circular exist in the archive, the system will retrieve whichever one is most similar to the question asked. It might be the current one. It might be the one from 2019.
The model isn’t making a reasoning error: it’s answering correctly based on what it was given. The error happens upstream, in the selection of the material.
This produces a particularly insidious effect. The system stays confident even when the source is outdated, because its confidence in the answer doesn’t come from verifying the content — it comes from how relevant the retrieval was. An inconsistent archive doesn’t produce visibly wrong answers: it produces plausible, occasionally outdated ones, which is worse.
There’s also an economic consequence. If every output has to be fully reverified because it isn’t clear which version it’s based on, the time saved disappears. The project stays technically functional and operationally useless.
What can be cleaned up during the project, and what can’t
The distinction is concrete, and it’s worth making early, because it determines how long and how costly the work will be.
Can be done during the project. Automatic deduplication of identical or near-identical documents. Optical character recognition with quality checks on scans that can be recovered. Automatic extraction of metadata: date, document type, originating office, regulatory references cited. Normalizing formats into a searchable structure. Automatically flagging documents with no date or references, which becomes a work list for the organization.
Can’t be done during the project, or isn’t worth doing. Deciding which version is current when that information doesn’t exist in any document: that’s a call that belongs to whoever is responsible for that process. Reconstructing the content of illegible scans. Formalizing tacit knowledge, which takes time from the people who have it and can’t be delegated to a tool. Reorganizing the organization’s entire document estate, which is a multi-year program with its own goals.
The dividing line is clear. Technology can sort, extract, link and flag. It can’t decide what’s correct when the organization hasn’t decided.
A realistic path
The most common mistake is sizing the project around the entire archive. It’s the most effective way to never reach production.
The path that works, in our experience, starts from a narrow, well-governed document domain.
You choose a limited corpus: one type of case file, one body of regulation, one product family. The selection criterion isn’t size, it’s governance: better a small archive with a clear owner than a large one with none.
You measure the starting point. How many documents, how much duplication, what share is machine-readable, how many have a certain date. These are numbers you can get in a few days, and they change the conversation, because they replace impressions with facts.
You clean up what’s automatable and bring what isn’t to the organization’s attention, as a specific list rather than a general observation.
You build a set of evaluation questions together with the people who will use the system, including the hard cases and the exceptions. That’s the basis for measuring accuracy in a verifiable way, instead of relying on how the first few uses feel.
You extend to other domains only after validation, reusing the cleanup rules already defined.
What this path doesn’t solve
It doesn’t produce a perfect archive, and that isn’t the goal.
It doesn’t replace a document management policy. If there’s no rule for who publishes, who updates and who retires a document, the disorder rebuilds itself within a few months.
It doesn’t remove the need for human oversight on the cases that matter, especially where the output feeds into a formal act or assessment.
And it doesn’t turn a poorly defined process into a clear one. Document analysis often makes it obvious that two offices follow different practices. Making that visible is useful. Deciding it is up to the organization.
In closing
Before choosing a model, it’s worth knowing what state the documents it will work on are actually in. That’s a check that takes a few weeks and, in most cases, changes the scope of the project and reduces its risk.
If you want an objective read on your document base before committing a budget, the assessment is where we start: measuring the current state, identifying the most promising domains, and estimating realistically how much preparation work is involved.




