A training pipeline is many processing locations

If your transfer assessment covers the inference endpoint and stops there, this guide maps the systems a training pipeline actually touches and shows where the exposure usually sits. The pattern is common: legal review looks at the API that serves predictions, signs its DPA, checks its hosting region, and closes the file. Meanwhile the fine-tuning run that preceded it copied the corpus through a preprocessing notebook, sent a slice to a labelling vendor, logged hyperparameters to a tracking service, pushed checkpoints to object storage, and pulled a base model from a registry. Each of those is its own processing location, and any one of them can put personal data outside the intended jurisdiction.

The endpoint gets the attention because it is the visible part. It has a URL, a contract and a dashboard. The training side is a chain of systems bought by different teams at different times, and Chapter V of the GDPR does not care how the chain is organised internally. The EDPB's Recommendations 01/2020 require exporters to know every transfer, including onward transfers made by processors acting on their instructions, so a sub-processor in a third country is your transfer, not your vendor's problem.

The EDPB sets out 3 cumulative criteria: an exporter subject to the GDPR, personal data made available to another controller or processor, and an importer in a third country or international organisation. A border crossing or foreign owner alone does not settle whether a disclosure is a Chapter V transfer.

  • Source data store: where the raw records live before anything happens to them
  • Preprocessing: the notebooks or jobs that filter, clean and split the corpus
  • Annotation or labelling: the vendor whose reviewers read the raw records
  • Training compute: the GPUs, their host, and the storage attached to them
  • Experiment tracking: metrics, configs and often sample data sent to a SaaS service
  • Artefact and checkpoint storage: object storage holding intermediate datasets and weights
  • Container registry: where the training image is pulled from
  • Evaluation: the harness, its datasets, and whoever can see the outputs

What a per-system assessment gives you

Running the assessment once for the pipeline produces one answer where you need eight. The EDPB's Recommendations 01/2020 define six steps, and each step can resolve differently per system: the tracking SaaS in the US and the GPU cluster in Spain do not fail for the same reasons, and they are not fixed by the same measures.

  1. Map each transfer, including onward disclosures and remote access
  2. Identify an applicable transfer mechanism for each recipient
  3. Where required, assess whether that mechanism remains effective under the destination country’s law and practice
  4. Adopt effective supplementary measures where needed; suspend or avoid a transfer if protection cannot be ensured
  5. Complete the procedures required for the chosen mechanism and document the assessment
  6. Reassess when the recipient, processing, law or access arrangements change

The US position is the worked example most pipelines need first. Since the Commission's adequacy decision of 10 July 2023, personal data can flow to US companies certified under the EU-US Data Privacy Framework without a further transfer tool, while recipients that are not certified still need one such as SCCs, backed where necessary by supplementary measures. That distinction is the direct outcome of Case C-311/18 (Schrems II, 16 July 2020), which invalidated Privacy Shield and required exporters to verify their transfer tools actually deliver protection in the destination jurisdiction. Practically: check the DPF list for the tracking vendor, but do not assume the annotation platform or the registry mirror is on it.

For the compute provider, the per-system answer is documentary. The DPA carries the named sub-processor list and is available on request, which is exactly the artefact step one and step two of this assessment need for the training compute entry: you can see who processes, where, and under which onward-transfer terms. For a deeper treatment of legal bases and data-subject rights on the training side, see our GDPR AI Training Data Processing: A Technical Compliance Guide.

Mapping every system the data touches

Map each system’s operator, data, location and remote-access paths. Apply all 3 EDPB criteria to each disclosure: a GDPR-subject exporter, personal data made available to another controller or processor, and an importer in a third country or international organisation. Access by a separate group company can qualify; a travelling employee of the same entity is a different case.

SystemTransfer question to answer
Source data storeWho operates it, in which region, and can vendor support staff in a third country open a bucket or a ticket with data access
PreprocessingWhere do the jobs run, and do they write intermediate outputs to a different region than the source
Annotation or labellingWhere do the human reviewers sit, and what sub-processors sit behind the platform
Training computeWhere do the GPUs and their attached storage physically run, and who can access the nodes remotely
Experiment trackingWhich SaaS receives configs, metrics and sample records, and where is it hosted
Artefact and checkpoint storageWhich region holds the buckets, and does replication or backup cross a border
Container registryWhere is the image hosted, and does it embed credentials or sample data in layers
EvaluationWhere do evaluation datasets live, and who reads the model outputs

For Serverless Training, map the compute and checkpoint storage separately until the actual locations and access arrangements are confirmed. An approved statement about a European fleet does not establish one shared site. Include the container registry when uploaded image layers, build logs or credentials disclose personal data. Downloading a public image does not itself send the training corpus to the registry.

Annotation vendors are the missed exposure

Annotation escapes the map for three structural reasons. A different team buys it, usually the data or product side, so the procurement review that covered the cloud contract never sees it. The work is human by definition, meaning raw records, including free-text fields nobody pseudonymised, are displayed to people outside your organisation. And the review workforce is often located outside the EU, because that is where the labelling capacity is priced.

Remote access can constitute a transfer when personal data is made available to another controller or processor in a third country and the other EDPB criteria are met. A separate annotation vendor displaying EU-hosted records is a common example. Access by an employee of the same entity is not automatically a Chapter V transfer, although foreign access can still create security risks that need assessment.

What to demand from the vendor is short and checkable:

  • Where the reviewers sit, per project, in writing, with a commitment to notify on workforce changes
  • The applicable transfer mechanism: for example, a covered US recipient’s Data Privacy Framework certification, another adequacy decision or appropriate safeguards such as standard contractual clauses
  • The sub-processor chain behind the platform, including the tooling vendors the reviewers' workflow runs on
  • Whether raw records are minimised or pseudonymised before they reach the labelling UI at all

If the vendor cannot answer the first question without an escalation cycle, treat that as the assessment result. The vendor vetting cluster on our magazine collects the provider-side version of this checklist, starting with The DPA Question: Sub-Processors in AI Inference.

Minimising before the data leaves the source

Minimise personal data before building the training corpus. Remove fields that the training or labelling task does not need, and assess whether identifiers can be pseudonymised without undermining the task. These measures can reduce disclosure, but they do not automatically make the remaining data anonymous or remove transfer duties.

  • Filter irrelevant personal data before the corpus is built, so records the task does not need never enter the pipeline
  • Drop free-text fields the task does not require, since free text is where identifiers hide
  • Pseudonymise identifiers at ingestion, replacing direct identifiers with tokens held separately
  • Restrict the annotation export to the minimum fields the labelling task actually displays

State the limit honestly: under Recital 26 of the GDPR, pseudonymised data that could be attributed to a natural person with additional information is still personal data, so these measures reduce exposure and shrink the population of affected data subjects, but they do not take the pipeline out of Chapter V. EDPB Opinion 28/2024 expects exactly this kind of preparation-stage filtering and mitigation to be considered and documented when personal data is used to develop AI models, which makes the ingestion decisions above part of the compliance record rather than an engineering preference.

Whether model weights carry the data forward

The question every fine-tuning team eventually asks is whether the trained weights are personal data. If they were anonymous, the transfer map could end at the checkpoint: weights could move to a third country for serving without a Chapter V analysis. The EDPB's answer in Opinion 28/2024 is that models trained with personal data cannot in all cases be considered anonymous, and whether a specific model is anonymous must be assessed case by case, looking at the likelihood of extracting training data from the model itself and of obtaining personal data through queries to it.

That is an unsettled question, not a resolved one, and it is worth saying plainly because a lot of internal documentation assumes the friendly answer. A fine-tuned model trained on a small, distinctive corpus is a different extraction surface from a foundation model trained on web-scale data, and the EDPB's case-by-case framing means your corpus size and memorisation exposure are part of the assessment. If you are also weighing what fine-tuning means for your AI Act obligations, that is a separate question covered in Does Fine-Tuning Make You a GPAI Provider?

The transfer assessment does not automatically end at the weights. Document extraction risk, corpus characteristics and access to model outputs. Keeping compute and serving in Europe can reduce some transfer paths, but does not settle whether weights contain personal data or meet the other GDPR duties.

Deleting the checkpoints and logs too

Storage limitation requires keeping personal data no longer than necessary for the stated purpose, subject to applicable exceptions. Finishing a training run does not automatically make every checkpoint or log unnecessary. Set justified retention periods for source data, intermediate files, checkpoints, logs and evaluation outputs, then document deletion or review at the end of each period.

  • Intermediate datasets written by preprocessing jobs, often in a different bucket than the source
  • Checkpoints from every epoch, which can each reproduce corpus fragments through the weights question above
  • Experiment tracking logs, which frequently embed sample rows for debugging
  • Evaluation outputs, which can contain model responses to real records

Platform behaviour matters here, because deletion you cannot execute is deletion you cannot document. On the platform, object storage files persist until the customer deletes them, so checkpoint and dataset deletion sits in your hands and belongs in your deletion runbook as an explicit step with an owner and a schedule. Write the runbook against the map from section three: every row that stores data gets a deletion entry, not just the source store.

Use the data-flow map to assign a justified retention period, access control and deletion owner to every stored copy. Confirm compute and storage locations separately with the provider. Serverless Training can simplify operations, but it does not remove the need to map external registries, annotation services, tracking and evaluation.