Skip to main content
Artificial Intelligence · Business & Lab

Where Should Sensitive AI Data Be Processed?

← Artificial Intelligence

Written by Evren BalPublished  · 9 min read

A data-flow plan shows backup and support copies beside a contract, with one copy under a magnifier.
Discuss this article with your AI

TL;DR — Keeping sensitive AI data on-site or within the EU/EEA is a sensible option to consider first. It can avoid some transfers, but only if support access, logs and backups stay within the same boundary. Choose a service by checking the whole data flow and deciding who will operate its safeguards. If the legal requirements cannot be met, narrow or stop the project; a stronger contract alone cannot fix that.

Consider a hypothetical European manufacturer. A service agent receives a call about a machine failure. To understand the case, they must pull together emails, service tickets, parts history, technician notes, and the prior customer conversation. A search system connected to a language model could gather the relevant records and draft a summary. Those records may contain names and contact details, but also confidential descriptions of faults, production interruptions, and disputes.

The company wants its service team to spend less time reconstructing what happened. Before it buys a tool, it needs to decide which records that tool should receive, who else could see them, and whether the team can operate it reliably.

My earlier article on national AI infrastructure examined Türkiye's EVREN platform. Its emphasis on local processing raises a question for companies that do not plan to use that platform: where should their own sensitive AI data go? For a European company, that may mean comparing an on-site system, a domestic provider and a service operating within the EU/EEA.

I have also examined what happens when chatbot conversations reach an external AI provider. Here the focus is the broader buying decision: which processing arrangement fits the data, the task and the team that must run it?

Sensitive data is not just a GDPR category

There are two reasons the manufacturer's records need protection.

Personal data is information relating to an identified or identifiable person. Under the GDPR, pseudonymised data can still be personal data; pseudonymisation is a safeguard, not a declaration that privacy rules no longer apply. The Regulation requires a lawful basis for processing, an additional condition for special categories such as health data, appropriate processor terms where a provider acts for the organisation, and security measures appropriate to risk. It also has a separate framework for transfers to third countries.

Commercially sensitive information can be damaging to expose even when it contains no personal data. A defect pattern, pricing dispute, supplier failure, source code fragment, or acquisition plan may fall into this category. The company may also want to prevent a provider from retaining it unnecessarily or using it for model training. In our example, a single service record could contain both a technician's contact details and a customer's confidential production information.

The sector alone does not tell you how sensitive a task is. Editing a clinic's public opening-hours notice is different from summarising a patient's medical history. A manufacturer's service notes could also mention an employee's injury. Start with the actual records, not a policy that treats every task in an industry alike.

Data residency is a map, not a label

“EU-hosted” is a useful starting point, but what does it cover? Check where the model runs, where the application stores its searchable records, where logs and backups go, and from which countries support staff can access them. The provider may use other companies to deliver parts of the service. That access belongs in the same assessment.

For the GDPR's rules on international transfers, the relevant boundary is the European Economic Area: the EU plus Norway, Iceland and Liechtenstein. The GDPR does not generally require a German company to keep all its data in Germany, or an Austrian company to keep it in Austria. Separate sector rules, confidentiality duties or contractual restrictions still need checking.

Suppose the manufacturer's model runs in an EU data centre, but the application writes the full input into an error log. If that log goes to a monitoring service outside the EEA, a copy of the service record goes there too. A backup may create another copy. A separate support provider outside the EEA may also be able to read the records remotely. Checking the model's address would miss all three.

The European Data Protection Board’s Guidelines 05/2021 illustrate the distinction carefully: a processor in a third country remotely accessing personal data held in the EU can create a transfer, while an employee of the same controller travelling abroad is not automatically the same kind of transfer. The details of the parties and access matter. Treating every overseas connection alike is too broad; treating a remote provider access path as irrelevant is too casual.

If processing, storage and the relevant provider access genuinely remain within the EEA, the company can avoid third-country transfers for those activities. That is a concrete reason to examine regional options first. It does not prove that access is well controlled, records are deleted on time, or the service is suitable for the job.

A promise not to train on customer data answers another useful question, but only that question. It does not tell you when the provider deletes a record or whether support staff can read it.

Start with the data path that the task actually needs

First narrow what the customer-history tool needs to do. It may need the machine model, service chronology, approved repair notes, and current case status. It may not need full email archives, contact details, internal legal commentary, or every attachment ever uploaded.

The company should also test whether a better service workspace would solve the problem without AI. A structured timeline, clearer case ownership and ordinary search might be enough. Where records already follow a consistent format, software can assemble them using fixed rules. The team may need better-organised records more than it needs generated summaries.

If an AI summary is still justified, start with synthetic records created for testing. Decide which fields the summary needs and which material must stay out. Removing names from real records is not proof of anonymisation; the remaining context may still identify someone. Real test data therefore needs its own assessment and permissions. Test the summary's accuracy as well: keeping a record private does not make a wrong repair history useful.

The German Data Protection Conference's 2024 guidance takes a similar approach: define the purpose, establish who is responsible, and make clear which uses employees are allowed to make of the system. It is practical guidance for organisations deploying AI, not a certification or an exhaustive list of requirements.

Three workable operating models

Location and operating responsibility are separate choices. A company can buy a managed service in its own country without buying servers. It can also rent local capacity and run the model itself. The main trade-offs are:

ArrangementWhat it can simplifyWhat your team still has to doWatch for
Managed AI serviceThe provider runs the model and underlying service.Assess provider access, retention and permitted uses; agree suitable terms; control the application's own data flows.Regional processing may coexist with overseas support or backups.
Rented EU/EEA infrastructure, with a model your team operatesYour team chooses the model and configures the application and logging.Secure and update the software, manage access, handle incidents and plan capacity; assess the infrastructure provider too.Your own logging or support choices can expose data.
On-premises or on-device processingDirect control over the equipment and network connections.Maintain equipment, security, recovery and model updates, with people able to support them.Remote support, updates and monitoring can still create external dependencies.

Buying a managed service reduces operational work, not the need to assess the provider. Running the system yourself gives your team more decisions to make, but it also needs the people to carry them out. Compare the options on output quality, response time, capacity and total operating cost, including maintenance and recovery. A system the company cannot reliably support is not a good choice merely because it is nearby.

When regional hosting is not enough—or not feasible

The preferred location may not offer a model that performs well enough, sufficient capacity or the support the company needs. An external service can be considered only if the applicable requirements allow it. For personal-data transfers outside the EEA, identify a valid route, such as an applicable adequacy decision or appropriate safeguards, including standard contractual clauses where suitable. A commercial service agreement is not a substitute for that assessment.

Encryption also needs a closer look. A service may encrypt data during transmission and storage, yet need to read it to perform the task. The EDPB's Recommendations 01/2020 address a specific case: where the processor needs readable data, holds the keys, and is subject to problematic third-country access laws, those two forms of encryption do not by themselves solve that access problem. This is not a general ban on cloud services. The protection must work under the actual conditions of the transfer.

Before signing, ask the provider to explain which companies can access the records, from where, and for what purpose. Agree retention periods, permitted uses and what happens to copies in backups. Where the provider is a processor, the data processing agreement must match the parties' roles. Establish how incidents will be reported and investigated, what evidence you can inspect, and how data will be returned or deleted when the service ends. Then check that the service's access controls, settings and records support those commitments.

If a binding location restriction applies, or the conditions for a necessary transfer cannot be met, choose another design or stop. If the provider cannot explain its access arrangements, do not resolve the uncertainty by uploading real records and hoping the contract covers them.

Make the next decision before using real records

For the manufacturer, the next step is to trace one proposed service record through the application with the service manager, technical owner, security, privacy and legal teams. Follow the original record, the material sent to the model, the summary, the error logs and the backups. For each copy, record where it goes, who can see it and when it will be deleted. Resolve the unknowns before introducing real customer histories.

Use synthetic cases to test whether the proposed tool helps the service team and whether the chosen protections work. The result may justify a managed service, a model the company operates itself, or a simpler workspace without AI. The decision should leave the company able to explain how it protects the records and who will keep the service working. If it cannot do that yet, the next step is to close those gaps, not to add more data.

This article offers a general decision framework, not legal advice. A specific deployment needs assessment based on its data, jurisdiction and sector requirements.

Further reading

If this article was useful

Linking to it from a relevant page on your website or sharing it on social media genuinely helps it reach more people. Thank you for your support.

Linking and brand guidelines →

About this article

Use of artificial intelligence
AI-assisted — Evren Bal determined the subject and intended argument. AI-assisted tools supported source research, drafting and language review, and generated the cover illustration. The text is awaiting Evren Bal's final review and release approval.