Futura AI
it
All articles
  • Architecture
  • Industry

Beyond text: where Speech AI and Vision AI solve real problems

Not all company information is written down. Call transcription, technical drawing reading and visual inspection in production: where these work today, and under what conditions.

by Daniele Grotti4 min readUpdated on
Futura AI — Beyond text: where Speech AI and Vision AI solve real problems

Not all company information is written down. Some of it is spoken, some of it is drawn, some of it is photographed on the shop floor.

Corporate AI projects almost always focus on documents, for a practical reason: they’re already digital and already archived. But some of the most operationally relevant information circulates in other forms and isn’t kept in a usable way.

Three areas deserve attention, each with its own requirements and limits.

Transcribing and analyzing support calls

A support service generates dozens or hundreds of conversations every day. What remains is the ticket record, filled in quickly at the end of the call, with a summary that varies from operator to operator.

The detailed information — how the customer described the problem and what checks were done — gets lost.

What’s practical today. Automatic transcription in Italian has reached a level that makes it usable under ordinary conditions. Built on top of transcription: ticket pre-filling, extraction of references mentioned, classification of the problem type, and aggregate monthly analysis of recurring causes.

The last point carries the most value and gets the least attention. Knowing which problems generate the most contacts, broken down by product and by period, is information that today doesn’t exist in structured form in many companies, and it says as much about product quality as about service efficiency.

The requirements. Sufficient audio quality, which isn’t a given on phone lines and in noisy environments. Handling overlapping speakers. And above all, compliance with data protection obligations and rules on remote monitoring of employees: notice, legal basis, involvement of employee representatives where required. These are obligations to address before the project starts, not during it.

Reading technical drawings, labels and poorly scanned documents

This is the area with the most direct return, because it acts on daily, measurable activities.

Technical drawings. Extracting dimensional callouts, tolerances and finish indications from drawings in non-uniform formats is now practical with models that interpret the drawing and the text together. On a manufacturing project, we measured 86 percent accuracy in automatically extracting dimensions from the client’s real CAD drawings, with low-confidence cases flagged for review by the technical office.

The figure also points to the right way to set up these systems: human review on uncertain cases is part of the architecture, not a fallback.

Labels and nameplates. Reading serial numbers, codes and nameplate data from photographs taken in the field. Useful in field service, where the technician photographs the nameplate instead of transcribing it, eliminating transcription errors on alphanumeric codes.

Low-quality scans. Current models handle historical documents, handwritten forms and photographs of pages better than traditional optical character recognition, because they use context to interpret uncertain characters. With a consequence to manage: by using context, they can produce a plausible but wrong reading. Critical documents need verification, and numeric values need to be handled with stricter thresholds than text.

Visual inspection in production

This is the most mature of the three areas and the most distant from the others, because it isn’t about generative AI but about computer vision, which has a well-established industrial history.

Surface inspection, checking the presence and position of components, dimensional conformity checks, in-line code reading.

The requirements are strict and non-negotiable. Controlled, consistent lighting. Repeatable positioning of the part. A set of images of both conforming and non-conforming examples — for rare defects, this is the main constraint: if a defect occurs only a few times a year, collecting enough examples takes time.

Evaluation needs to be set up correctly. In a quality check, the two types of error aren’t equivalent. Rejecting a good part costs that part; approving a defective one can cost far more. Calibration has to be done on this basis and agreed with the quality function, not optimized for overall accuracy.

And there’s still the question of validation. Where the check is relevant to product safety or to compliance, the system runs alongside the existing check and doesn’t replace it without formal validation.

Current limits and the quality of the input data

One thing is common to all three areas: the quality of the input data matters more here than it does in text processing.

A poorly formatted text document remains readable. Noisy audio, a blurry photograph or variable lighting degrade the result directly, and it can’t be recovered downstream.

The consequence is that a significant part of these projects is about acquisition conditions: defining minimum requirements, equipping operators with the right tools, standardizing capture procedures. It’s unglamorous work, and it determines the outcome.

Three further limits need to be stated.

Transcription degrades noticeably with strong accents, highly specialized terminology and overlapping voices.

Interpreting technical drawings works on the most common conventions; specific symbologies or internal company standards require dedicated configuration.

Visual inspection for rare defects remains an open problem, due to the lack of available examples.

In closing

Non-textual information is often the information closest to day-to-day operations, and for that reason the least structured. Making it usable produces measurable benefits, provided acquisition conditions are addressed first.

The selection criterion is the same as for the other areas: a narrow scope, defined quality requirements, a baseline measurement, and a confidence threshold below which the system flags instead of deciding.

If you have a flow of non-textual information that’s currently lost or handled by hand, we’re available for a conversation about that case.

If this topic touches a real process in your organization, let's talk about it with a focused AI Assessment.

Request an AI Assessment