Measured results
Every figure published on this site links to the case study that states its scope, sample and measurement method.
This page brings together, in one place, the metrics of the five systems in production with clients who authorized being named. It adds no new numbers: it shows the scope and method next to each figure, so you don't need to open the full case study to check them.
Industry
Automated extraction of dimensions from CAD drawings and ISO conformity checks
86% accuracy extracting dimensions from CAD drawings; quote preparation goes from 4-5 person-hours to a few minutes.
Scope and measurement method
- Scope: the subset of most frequent technical drawings and the ISO standards applicable to the product families selected for the rollout.
- Evaluation set: a sample of 74 technical drawings from the client, totaling 3,842 dimensional callouts — assemblies, machined parts and drawings with varying levels of geometric and annotation complexity — with dimensions read and manually verified by the engineering office as the reference for comparison.
- Metric: dimensional callouts extracted correctly out of the total callouts present in the sample. The denominator also includes the callouts the system did not read, not only those it produced an output for.
- Low-confidence cases: flagged for engineering-office review and counted as not extracted, not excluded from the calculation.
- The remaining 14% is why human review of uncertain cases is part of the architecture rather than a fallback.
Finance
Automated reconciliation of incoming payments: from the payment reference to the payer's NDG
92% of transactions reconciled automatically on a sample of roughly 15,000 verified transactions.
Scope and measurement method
- Scope: incoming payment flows in the order of hundreds of thousands of transactions a year, on the most recurring reference types, progressively extended to less regular structures.
- Evaluation set: a sample of roughly 15,000 transactions with the correct NDG already attributed manually by operators, used as the reference for comparison.
- What is measured: the share of transactions reconciled automatically out of the total, and among those the share attributed to the correct NDG. The two have to be read together: raising the threshold improves the second and worsens the first.
- The two errors are counted separately. A payment left to the operator is a processing cost; a payment attributed to the wrong position is an error that propagates, and weighs more.
- Measured share: 92% of transactions reconciled automatically on a sample of roughly 15,000 transactions, in the January–March 2026 period. On the same sample, the wrong-attribution rate is 0.8%.
Technology and AI
Inbound voice agent with answers grounded in a knowledge base
90% of in-scope calls completed with no human intervention; low-confidence cases go to an operator.
Scope and measurement method
- Scope: a defined subset of recurring requests, with out-of-scope requests routed to an operator by design and not as a fallback.
- Evaluation set: real conversations reviewed and scored, including the calls where the agent correctly declined to answer.
- What is measured: correctness of the answers given, share of calls completed without human intervention, and escalation rate. The last is not a defect to minimize: a well-judged escalation is worth more than a risky answer.
- In voice, abstention matters more than in text: the listener cannot check the source while talking, and a wrong answer delivered fluently offers nothing to catch it on.
- Measured share: 90% of calls completed without human intervention, based on the real conversations reviewed and scored. Escalation rate: 70%, calculated on the low-confidence cases only (the remaining 10% of calls), not on the total. Observation period: March–May 2025.
Finance
Semantic search over regulatory and compliance documentation
86% accuracy in classification and extraction over roughly 25 million pages; no single headline percentage for the RAG itself.
Scope and measurement method
- Classification and extraction phase (DWH): 86% accuracy on a corpus of roughly 25 million pages, measured before the semantic search system went live, as the condition for building the RAG on normalized data.
- Scope of the semantic search (RAG): pilot phase with a single compliance team, extended to other departments only after validation on accuracy, sources and behavior on ambiguous cases.
- Evaluation set: 186 questions built together with the compliance team — 112 factual/documentary questions, 48 applied and procedural scenarios, 26 deliberately ambiguous or "adversarial" questions — including the hard cases and the questions the archive holds no answer to.
- What is measured on the RAG: accuracy on the answers given, coverage over the questions asked, and citation correctness, all assessed on the same question set.
- We do not state a single headline percentage for the semantic search: three quantities matter, they move in opposite directions as thresholds change, and reporting only one would be misleading. The upstream classification phase is different: accuracy there is a single, well-defined metric — hence the 86% above.
Technology and product
A RAG-based music recommendation engine (CORRD)
85% of recommendations judged relevant by real users; 65% catalog coverage.
Scope and measurement method
- Scope: the subset of the catalog indexed at the start, progressively extended to the rest.
- Evaluation set: recommendations judged by real people, not offline metrics alone.
- What is monitored: relevance of the suggestion, diversity, and catalog coverage. They have to be read together because they move in opposite directions: an engine maximizing relevance ends up recommending the same things over and over.
- Recommendation quality is not accuracy: there is no correct answer to compare against. That is why evaluation needs people, and an evaluation set alone cannot tell you whether the engine works.
- Measured relevance: 85% of recommendations judged relevant by users in the evaluated sample. Measured coverage: 65% of the catalog reached compared with the previous recommendation logic. Catalog diversity remains tracked separately and has no published figure yet: high relevance achieved by always recommending the same titles would not be progress.
Have a similar process to measure?
The same method — baseline, agreed KPIs, post-release measurement — applied to your case. An AI Assessment establishes whether, where and how to act.
Request an AI Assessment