If an AI project doesn’t have an indicator defined before it starts, it won’t have a demonstrable result at the end.
Not because the result is missing. Because there won’t be a basis for comparison. Without an initial measurement, every later assessment becomes a comparison between the current situation and people’s memory of how work used to be done.
It’s the most frequent mistake and the most costly one, because it shows up at the worst possible moment: when leadership asks whether it’s worth continuing.
Three families of indicators
Almost every project of this kind creates value in at least one of three directions. It’s worth choosing one as the primary indicator and treating the others as secondary.
Time. How long an activity takes from the moment it enters to the moment it exits. Average processing time for a case, time spent searching for a piece of technical information, time to prepare a quote.
It’s the easiest family to measure and the quickest to communicate. The risk is measuring the time of a single step instead of the process: a faster step that creates a queue downstream produces no benefit.
Error. How often the result has to be redone. Cases sent back for incomplete documentation, quotes corrected after being sent, customer responses that need amending, audit findings.
It’s the most underrated family, and often the one with the greatest economic value, because an error costs the original work plus the rework plus, sometimes, the relationship with the other party.
Capacity. How much work the organization can handle with the same number of people. Cases closed per month, proposals submitted against requests received, backlog absorbed.
It’s the most relevant family when the constraint is volume rather than cost: agencies with a growing backlog, technical departments that give up responding to requests for proposals for lack of time.
How to build the baseline
The initial measurement has to be done before development starts, and it takes less time than people think.
Define the unit of measure unambiguously. “Processing time” needs a stated start point and end point. If the data can’t be pulled from the systems, sample it over a representative period.
Measure on a sufficient and honest sample. Including the hard cases, not just the ordinary ones. A baseline built on easy cases produces a flattering comparison with no value.
Record the variability, not just the average. Knowing the average time is six days tells you little if ninety percent of cases close in two and the rest in thirty. In many processes, the value of automation lies precisely in cutting down the long tail.
Note the context: available staff, seasonality, exceptional events during the measurement period. It’s needed to defend the comparison when someone, legitimately, questions it.
The most common mistake: measuring the tool instead of the process
Many project dashboards report the number of active users, queries per day, sessions per user, user satisfaction ratings.
These are adoption indicators. They’re useful for understanding whether the tool is being used and for catching organizational rejection early. They say nothing about return.
A heavily used system may have no effect on the process at all, if people consult it in addition to what they used to do instead of in place of it. It’s a situation we see fairly often in the first months, and it gets fixed by intervening on the workflow, not on the tool.
The reverse rule is a useful check: if the chosen indicator could improve even if nobody used the system, it’s not the right indicator. And if it could stay exactly the same while the system is used intensively, it isn’t either.
A calculation model, with an illustrative example
The base calculation requires four quantities: annual volume, time saved per unit, fully loaded hourly cost, actual adoption rate.
Gross annual benefit = volume × time saved × hourly cost × adoption.
From this you subtract the total cost of ownership: development, integration, licenses and usage costs, internal oversight, ongoing maintenance.
A numerical example, built to show the method and not tied to a real project. An office handles 6,000 cases a year. Document review and summary preparation currently takes 40 minutes per case. The system reduces that step to 15 minutes on standard cases, which are 70 percent of the total. Actual adoption in the first year is 80 percent.
Time saved: 6,000 × 0.70 × 0.80 × 25 minutes, roughly 1,400 hours a year. At a fully loaded hourly cost of 35 euros, the gross benefit is about 49,000 euros a year, from which the total cost of ownership is subtracted to get the net return.
Three warnings about this kind of calculation. The adoption rate should be estimated conservatively, because it’s the variable most often overestimated. The freed-up hours are only a benefit if they’re redirected to valuable activities, and that’s an organizational fact, not a technical one. And the benefits from the error family, often the most significant, don’t enter this calculation: they need to be assessed separately.
What measurement doesn’t capture
Some real effects don’t fit into a model like this.
The reduction in operational risk. A more systematic completeness check lowers the probability of an audit finding. It’s a probabilistic benefit, hard to attribute to a single year.
The continuity of knowledge when an experienced person leaves the organization. It has clear value and no market price.
The perceived quality of service toward citizens and customers, which is measured with different tools and over longer horizons.
It’s also worth remembering that the return is never immediate. The first months involve hand-holding, corrections and natural distrust. A model that promises full benefits from the first quarter is a model that will be proven wrong.
In closing
Measurement isn’t an end-of-project formality. It’s an initial choice that determines whether the project will be assessable at all.
You need a primary indicator, an honest initial measurement and an agreed review cadence. Three elements that take a few weeks to define and that change the quality of every later discussion with leadership.
In our assessment, defining the indicators comes before the technical choices. If you’re evaluating a project and want to understand what numbers it would be measurable against, that’s where we start.




