Futura AI
it
All articles
  • Data
  • Industry

How we measure 86% accuracy extracting data from CAD drawings

86% accuracy isn't a number you just state: it's a scope, a sample and a metric that can be verified. The method behind the published figure for the Gruppo SAG case.

by Daniele Grotti4 min read
Futura AI — How we measure 86% accuracy extracting data from CAD drawings

Eighty-six percent accuracy is a figure that’s easy to put on a slide and hard to verify, unless you also say what it was calculated on.

This article works backward: it starts from the number published in the Gruppo SAG case and shows the method behind it. Not to add technical detail for its own sake, but because a number without a scope, a sample and a metric isn’t a measurement — it’s a promise.

The problem the number describes

Gruppo SAG receives technical CAD drawings from different clients and suppliers, in inconsistent formats and graphic standards. Reading dimensional callouts, tolerances, and checking conformity against the applicable ISO, EN and DIN standards used to be done by hand by the engineering office: slow, repetitive, exposed to transcription errors, with a direct impact on quoting and quality.

The system built reads the drawings with Vision-Language models and OCR, extracts dimensions and tolerances, compares them against the relevant standards, and flags deviations and low-confidence cases for review. It’s the same pattern described, as a scenario, in the case study on extracting data from non-standardized documents: here, though, we’re talking about the real project, with the measured figure.

What we measured, exactly

Four elements, in the order they should be stated before you publish a percentage.

The scope. Not Gruppo SAG’s entire drawing archive, but the subset of the most recurring drawing types and the ISO standards applicable to the product families selected for the rollout. A percentage calculated on a different scope than the one stated isn’t comparable to anything.

The comparison sample. A set of real technical drawings from the client, with dimensions read and manually verified by the engineering office as the reference. Not synthetic drawings, not cases picked to be favorable: the same material the engineering office works with every day.

The metric. Dimensional callouts extracted correctly out of the total callouts present in the sample. The denominator also includes the callouts the system didn’t read, not only the ones it produced an output for: it’s the difference between accuracy and coverage described in how to evaluate an AI system before production, and here the two are kept deliberately distinct.

Low-confidence cases. Flagged for engineering-office review and counted as not extracted, not excluded from the calculation. Excluding them would have artificially raised the percentage without changing the actual work still to be done.

The missing 14%

One dimensional callout in seven, in the measured sample, isn’t read correctly.

That’s not a detail to minimize: it’s the reason human review of uncertain cases is part of the architecture, not a fallback added out of caution. The system flags low-confidence cases instead of guessing, and those cases remain the engineering office’s workload until the reading is verified. It’s the same principle applied in the data-extraction project cited above: a narrow, stated scope produces a reliable system; an optimistic scope produces a system that works half as well on everything.

What we still don’t say

The number of drawings and callouts in the comparison sample isn’t public yet: it’s a figure to confirm with the client before making it explicit. Until it is, we avoid following the 86% with a context figure we haven’t verified in writing — a scope described in words is better than a second number that’s as unverifiable as the first.

Likewise, the measurement covers the subset of most recurring drawings and standards selected for the rollout, not Gruppo SAG’s entire drawing archive: extending it to other product families is the next phase of the project, not an already-achieved result.

Why we publish the method, not just the number

Because an organization that has to authorize a system like this on its own data doesn’t need another reassuring claim. It needs to know how to verify a claim — its own, before ours.

Gruppo SAG’s 86% and the quote preparation time, down from 4-5 person-hours to a few minutes, remain the two numbers we publish with the method behind them. They’re also a concrete example of what a Document Intelligence system does when the problem is reading documents that don’t follow a standard.

If you’re evaluating a similar project, we’re available to talk through the measurement method applicable to your case: which sample, which metric, which threshold would be defensible under scrutiny.

If this topic touches a real process in your organization, evaluate it with a focused AI Assessment.

Request an AI Assessment