Over the past two years, almost every organization we talk to has seen at least one generative AI demonstration. Many have commissioned one internally. Few today have a system actually in production, used every day by the people working the process.
It’s rarely a budget question. It’s almost never a model question. It’s a question of what gets asked of the system the moment it stops being a proof of concept.
What makes a demo easy?
A demo operates under chosen conditions.
The documents are curated: readable, up to date, consistent with each other. The questions are the ones the system knows how to answer, because they’ve been tested beforehand. There are no permissions to respect, because no one is accessing real data belonging to citizens, customers or employees.
Above all, no one has to answer for a mistake. If the demo gets it wrong, you rephrase the question and move on.
This doesn’t make demos useless. They’re good for verifying that a technical direction is viable and for building internal buy-in. The problem starts when they’re read as an estimate of the work required.
What does it take to bring an AI project into production?
Moving into production changes five things at once.
The data becomes real: historical scans, multiple versions of the same document, inconsistent formats, attachments nobody has ever classified.
Permissions become binding. A system that answers can’t show a user content that user couldn’t have opened in the source repository.
Traceability is required. Which source generated that answer, who performed that action, when, and with what outcome.
The load becomes continuous: not a demo session, but queries spread across the day, with response times that have to stay acceptable even at peak load.
And maintenance is required. Documents change, procedures change, vendors update or retire models.
Where do AI projects stall before production?
In our experience, projects that don’t make it past the pilot phase almost always get stuck on one of these four points.
1. The quality of the document base. An archive with duplicates, misaligned versions and low-quality scans produces inconsistent answers even with the best available model. Checking this isn’t a technical detail: it’s the first activity of the project. In an agency that processes applications, for example, the same circular can exist in three versions in three different folders, and none of the three makes clear which one is current.
2. Integration with the systems already in use. An assistant that doesn’t read from the protocol system, the ERP or the case-management system forces people to do the work twice. The value doesn’t come from the conversation: it comes from the fact that the system sees the same data the operator sees. This is where the project runs into the least visible and most expensive issues: available or missing APIs, permission management, master-data alignment.
3. Internal ownership. A system in production needs someone who answers for it: who approves changes, who checks outputs, who decides when to stop it. When the project stays an initiative with no operational owner, it survives for as long as leadership attention lasts. Then it fades out.
4. The absence of indicators. Without a measure defined before launch, evaluating the system becomes a matter of impression. And the impression of someone who saw the demo differs from that of someone who uses the tool every day on hard cases. Average processing time, share of cases handled without rework, reduction in completeness errors: these are numbers that need to be agreed on at the start, while it’s still possible to measure the starting point.
How do you set up a project so it reaches production?
There’s no shortcut, but there is a sequence that reduces the risk.
Start from a narrow, well-governed document domain. One type of case file, one product family, one body of regulation. A small perimeter makes it possible to verify accuracy on real cases and correct course while it’s still cheap to do so.
Define the indicators before you build. If you can’t say which number is supposed to improve, the project isn’t ready to start yet.
Design the stopping points alongside the features. Where the system has to ask for confirmation, when it has to state that it doesn’t know, which actions remain exclusively human. These constraints aren’t a limitation of the system: they’re what makes it usable in a context with formal accountability.
Release in phases, starting from a low-risk case, and extend only after operators have validated it.
And treat training as part of the project. The people using the system need to know its limits as well as its capabilities, or the first mistake becomes a reason to abandon it.
What this approach doesn’t solve
A structured approach doesn’t eliminate organizational problems.
If the document archive has no governance, AI will make that obvious, not fix it. If two offices don’t agree on which procedure is correct, no system will decide for them. If there’s no internal owner, the project will remain an experiment even if it’s technically successful.
The reverse is also true: some processes don’t need artificial intelligence. A deterministic rule or a workflow review can produce more value, with less risk and less maintenance cost. Saying so is part of the job.
In closing
The useful question at the start of a project isn’t whether AI works. It does.
The question is whether the organization has the data, the integration, the roles and the indicators to sustain its daily use. That’s an assessment that can be done in a few weeks, before committing a significant budget.
If you’re weighing a concrete case in your organization, we’re available for a technical conversation about that specific process: what data is needed, where the friction points are, and what would be realistic to measure in the first few months.




