Futura AI
it
All articles
  • Security
  • Compliance

Data security in AI projects: seven questions to ask your vendor

In organizations with high compliance requirements, security isn't a contract clause: it's a selection criterion to verify upfront, in writing.

by Daniele Grotti5 min readUpdated on
Futura AI — Data security in AI projects: seven questions to ask your vendor

In an organization with high compliance requirements, security isn’t a section of the contract: it’s a selection criterion.

The difference is operational. If security is a section of the contract, it gets addressed after the vendor has been chosen, and at that point clauses get negotiated on top of an architecture that’s already been decided. If it’s a selection criterion, some answers rule out a vendor before technical evaluation even starts.

The seven questions below need to be asked in writing, and asked again in writing whenever the answer is generic.

1. Where does the data reside, and who can access it

The question needs to be broken down. Where does it reside during processing, where does it reside at rest, where do backup copies reside, where do the logs reside.

These are four locations that can differ. A service with data at rest in the European region can process or log elsewhere.

On access: which categories of the vendor’s staff can access the data, under what circumstances, with what authorization and what logging. The correct answer isn’t “no one” — implausible for a service that requires maintenance — but a description of the procedure.

The subcontractor chain needs to be included too. A vendor that delivers the service by relying on third-party infrastructure passes on to you the constraints of those third parties, and the list needs to be available and current.

2. Is the data used to train models

Asked this way, the question almost always gets a reassuring answer. It needs to be broken down into three levels.

Is the data used to train or improve the vendor’s own models. Is the data passed on to third-party model providers, and under what contractual terms. Is the data retained for security, abuse-monitoring or service-improvement purposes, and for how long.

The third is the most overlooked. Several services retain interactions for a defined period even when they don’t use them for training, and in a context with sensitive data this retention needs to be known and assessed.

It also needs to be checked whether exclusion from training is the default configuration or an option to be activated, and whether it can be changed unilaterally through an update to the terms of service.

3. How are document permissions managed

This is the question that separates a system designed for regulated contexts from one merely adapted to them.

Three checks.

Does the permission filter apply at document retrieval or at answer presentation. It has to apply at retrieval: an unauthorized document must never enter the material an answer is built from.

Are permissions dynamically inherited from the source systems or statically replicated. A static replica drifts out of alignment, and the drift isn’t visible until it causes an incident.

How are changes handled: revoking access, a role change, termination of a relationship. And how quickly does the change take effect on the system.

4. What gets recorded in the logs

A system in a supervised context needs to log queries, documents retrieved, answers produced, actions performed, and the identity of the user who originated them.

The questions to ask cover log content, their location, retention period, who can consult them and through what procedure, and whether they’re exportable in a format usable for an internal audit.

One specific point: logs often contain personal or confidential data, because they record the content of queries. Their protection needs the same level of attention as the primary data, and that needs to be checked separately.

5. How is a piece of data deleted

Deletion needs to be verified across four layers: the indexed document, its representation in the search index, backup copies, and logs that contain extracts of it.

Three questions: which procedure, on what timeline, with what documentary evidence that deletion actually occurred.

The evidence is the critical point. If a right to erasure is exercised, the organization has to be able to demonstrate that the request was carried out, not merely forwarded to the vendor.

6. What happens if the service is interrupted

Two distinct scenarios, often conflated.

Temporary unavailability: what service levels are guaranteed, how are they measured, what are the contractual consequences, and what alternative way of working is provided for the period of unavailability. The last question is the most concrete and the least often asked.

Termination: what notice period is guaranteed if the service or a model is discontinued, what data and configurations are returned, in what format, within what timeframe, and with what migration support.

The scenario of the vendor ceasing operations needs to be included too. Escrow of configurations with a third party, or code availability under certain events, are standard clauses in contracts of this kind.

7. Who is accountable in the event of an incident

The last question is the one that makes all the previous ones verifiable.

What notification procedure, on what timeline. Who conducts the analysis. What obligations of cooperation apply in the event of a report to the supervisory authority, given the tight deadlines the law requires. How liability is split between vendor and client, and with what limitations.

The liability-limitation regime in particular needs to be checked. A cap tied to annual fees can be irrelevant against the impact of an incident involving sensitive data, and that’s a factor to weigh, not necessarily a reason for exclusion.

What these questions don’t cover

They don’t cover the application-level security of the system. Defense against prompt injection, control over executable actions, isolation of instructions, adversarial testing: these are separate topics, matters of design, and need their own set of questions.

They don’t cover the quality of the answers, which is measured through evaluation on real cases.

They don’t replace a data protection impact assessment where one is required, nor sector-specific compliance checks.

And they guarantee nothing if the answers aren’t put into the contract. A statement made during negotiation doesn’t carry the same weight as a clause.

In closing

A vendor who answers these seven questions precisely, in writing and without generic formulas, has already demonstrated something relevant about how they work.

The answers need to be collected before the technical evaluation, because some of them narrow the field decisively.

If you’re evaluating a project involving sensitive data, we’re available for a conversation about security requirements and data perimeter applied to your context.

If this topic touches a real process in your organization, let's talk about it with a focused AI Assessment.

Request an AI Assessment