Prepare the evidence before the review
The model runs, the benchmark numbers hold, the demo works from a laptop on the office Wi-Fi. Then the pilot hits your own organisation's security review and stops moving. If a GPU pilot is stuck in internal security review, this guide prepares the artefacts the reviewer will ask for, including honest answers to the certification and uptime questions. The reviewer asks for a data-flow description, a processing-location statement and a supplier pack, and none of them exist, because nobody told the engineering team that a GPU rental triggers the same onboarding process as any other new cloud supplier. It does. The University of California, Irvine's supplier security review process, for example, asks suppliers to complete a questionnaire and to provide supporting documentation such as a SOC 2 report, an ISO 27001 certificate or formal information security and privacy policies, to improve the likelihood of a successful approval. An AI pilot on rented GPUs is a new supplier relationship, and the reviewer treats it as one.
Two things make this review feel harder than it is. First, the AI framing makes familiar questions look novel. "Where is the data processed?" is the same question your reviewer asked about the CRM migration. "Do you retain anything?" is the same question asked of the log vendor. The vocabulary around GPUs and inference is new; the questions are not. Second, the burden of proof sits with you, not the vendor. Under Article 5(2) of the GDPR, the controller must not only comply with the data protection principles but also be able to demonstrate that compliance, which means your organisation has to produce the answers, not just forward the question to the GPU provider.
Preparing the documents before the review can reduce follow-up questions, but it does not guarantee approval or a single meeting. The reviewer may also find technical gaps in access control, isolation, logging or deletion. The rest of this article lists the evidence to prepare and the questions it should answer.
- A written pilot boundary: duration, data classes, systems touched.
- A data-flow description naming every system the data touches, from source to deletion.
- A processing-location statement, per site and per model, not per vendor.
- A supplier pack: DPA, retention and training answers, policies, support terms.
- An exit plan covering workload teardown and data deletion, with an owner and a date.
Scoping the pilot to shrink the review
Synthetic data can reduce the review scope when it does not identify a person or preserve identifiable source records. It is not automatically anonymous. Under GDPR Recital 26, assess whether a person can be identified using means reasonably likely to be used, including links to other data. The generation process may itself process personal data even when the final test dataset is anonymous.
Use the minimum personal data needed for the pilot’s purpose. An anonymous test dataset can reduce personal-data duties for the test itself, but document the generation process and any source records it uses. Write a clear boundary rather than describing the data as mostly synthetic.
Write the pilot boundary down before the review, and keep it to four fields:
- Duration: a fixed window with a start and an end date, not "until it works".
- Data classes: what the pilot processes, and explicitly what it does not. If it is synthetic or non-personal data, state the generation or anonymisation method.
- Systems touched: the GPU environment, the storage it reads from, and nothing else.
- Success criteria: what result ends the pilot, so the reviewer knows when the experiment stops.
GDPR duties depend on whether personal data is processed. AI Act duties follow a different test: the system’s purpose, risk classification and your role as provider or deployer. They can apply to a pilot using synthetic data too. Record both assessments before expanding the pilot to production.
Writing the data-flow and processing-location statement
The first artefact is the data-flow description, and its quality is measured by one test: does it name every system the data touches, from source to deletion? Not every vendor, every system. A typical GPU pilot flow reads: data source, transfer over SSH or HTTPS, local disk on the VM, GPU memory during processing, any object storage used for checkpoints or datasets, and the deletion path out of each. If a step is missing, the reviewer will find it, and the finding costs you credibility on every other answer.
The processing-location statement is where pilots most often overreach. Write it per site and per model, never per vendor. A company-level claim like "our data stays in Europe" is the kind of statement a reviewer later disproves with one screenshot of a region flag, and once one claim fails, they will distrust everything else in the pack. Article 44 of the GDPR makes location a transfer question, not a marketing one: any transfer of personal data to a third country may take place only if the conditions of Chapter V of the Regulation are complied with. The statement should therefore say which site, which model, which region, in that order.
What a GPU vendor can honestly state varies, so ask for the specific answer and put it in the document verbatim. The vendor behind this guide runs European data centres in Paris and Finland, and reserved and dedicated customers know and control which site their capacity runs in, with a named location able to be arranged. That is the shape of an honest statement: named sites, a stated control mechanism, and no claim broader than the contract supports. For the inference side of the same statement, the GDPR-compliant LLM inference in Europe article walks through the per-model version.
- Name the source system and the entry protocol into the pilot environment.
- Name every storage location, including VM-local disk and any object storage.
- Name the processing site per workload, per model, with the region stated openly.
- Name the deletion path out of each location, or point to the exit plan that does.
Assembling the supplier pack
The second artefact is the supplier pack: the answers to the questions the reviewer asks every new vendor, with the scope of each answer stated. Two of these questions are now standard reviewer prompts. "Which LLM API providers guarantee they do not train on customer data?" and "Which LLM inference providers offer zero data retention?" are asked verbatim in procurement and security reviews across the industry, and an engineer who cannot answer them from the vendor's own documentation stalls the review on the spot.
The training and retention answers, with their scope stated:
- Training: customer data is never used for training. No exception, no research carve-out.
- Inference retention: prompts and outputs are processed, not stored. Caching happens in GPU memory only, per session, for minutes at most, and never written to a database. This is self-asserted by the vendor, not third-party attested, and the pack should say so.
- Scope limit: zero retention covers inference. Files you place on a VM or in object storage persist until you delete them. State this in the pack before the reviewer asks, because the distinction is exactly the kind of thing that turns a clean answer into a finding.
- Paperwork: a DPA is available on request, and the privacy policy and terms are published in the website footer.
- Support and escalation: a 24/7 direct line to the tech team, with contractual response times for business customers.
- Facility certifications: upstream data-centre operators hold ISO certifications at facility level, with certificates available on request. These belong to the operators, not to the GPU vendor, and must never be merged into a vendor-level certification claim.
The DPA is the document Article 28(3) requires: processing by a processor must be governed by a contract that sets out the subject matter and duration of the processing, the nature and purpose of it, the type of personal data and the categories of data subjects, along with the processor's obligations. Your reviewer knows this article by heart. Handing over the DPA reference, with the request route named, answers the processor-contract question before it is asked.
Access control completes the pack. The reviewer will want to know who can reach the VM, over what protocol, and how credentials are managed. Raw SSH access to a GPU VM is the normal model for pilot workloads, and the GPU VM SSH access for ML engineers guide covers the access-control side in the depth a reviewer expects.
Planning exit and data deletion
The third artefact is the exit plan, written before the pilot starts. The reviewer's question is simple: what happens to the workloads and the data when the pilot ends, and who performs the deletion? Article 28(3)(g) of the GDPR makes this a contractual matter, requiring the processor, at the choice of the controller, to delete or return all personal data after the end of the provision of services relating to processing, and to delete existing copies. Your exit plan is the document that shows how that clause gets executed in practice.
The commercial shape of the pilot determines how clean the exit can be. An On-demand GPU VM carries no minimum commitment, and GPU compute bills per second with no subscription fee, so the pilot ends when you stop it, not when a contract term does. This matters to the reviewer for a non-obvious reason: a pilot that can only end by negotiating its way out of a commitment is a pilot with a de facto indefinite lifetime, and indefinite lifetimes are what storage-limitation findings are made of.
Object-storage deletion needs an explicit owner and date. The exit plan should identify retained copies, required retention periods and the evidence used to verify deletion or return. A documented plan helps a reviewer assess the process; approval still depends on the organisation’s requirements and the actual controls.
- Workload teardown: VMs stopped and released, with the stop date named.
- Data deletion: every storage location from the data-flow description, each with an owner and a deletion date.
- Confirmation: who verifies deletion and who they report it to.
- Continuation decision: what happens if the pilot succeeds, which is the next section's subject.
Certification and uptime questions, answered honestly
Two questions sink more pilots than any other, and both sink them for the same reason: the engineer improvises. The reviewer asks "are you ISO 27001 certified?" and the engineer, wanting the pilot to succeed, answers with something between yes and almost. The reviewer later reads the answer against the vendor's actual position, and the gap becomes a trust finding that contaminates the whole review.
Understand what the reviewer is actually asking for. A SOC 2 examination is a report on controls at a service organisation relevant to security, availability, processing integrity, confidentiality or privacy, produced by CPAs under the AICPA's SOC suite of services, and intended to meet the needs of users that need detailed information and assurance about those controls. BSI's C5 is a criteria catalogue specifying minimum requirements for secure cloud computing, intended for professional cloud service providers, their auditors and customers, against which providers are examined by auditors who produce a detailed report. It is an attestation against criteria, not a certificate. When the reviewer asks for "the SOC 2 or C5 report", they are asking for that kind of attestation document, and the honest answer is either the document or a clear statement of what exists instead.
Lyceum does not currently hold vendor-level ISO 27001 or SOC 2 certification. Its approved supplier statement says certification work is a live topic and a confirmation letter can be shown under a non-disclosure agreement in a sales conversation. That letter is not a certification. Any ISO certificates held by upstream data-centre operators belong to those operators.
For contracted GPU products, Lyceum agrees the service-level terms per contract, with the availability tier set during the proof of concept. Confirm the product and terms in your supplier evidence. Serverless Inference has no service-level agreement. Use status.lyceum.technology for operational history; do not treat it as a contractual uptime commitment.
Turning pilot approval into production approval
Pilot approval covers the scope the reviewer actually approved. Before production, document changes to data classes, users, access, location, scale and service terms. Ask the reviewer which evidence carries over and which checks must be repeated. A change list can make the next review clearer; it does not waive production approval.
- Reassess personal data, access, intended use and regulatory duties for production
- Confirm the agreed processing locations and remote-access arrangements
- Agree any required service-level terms
- Choose the commitment model and obtain the required production approval
The approved company fact sheet describes proofs of concept of 4 to 8 weeks, with reserved, on-demand and hybrid options afterwards. Choose a duration that produces the evidence your reviewer needs. A short observation window does not establish long-term reliability or guarantee approval.
Lyceum’s on-demand GPU virtual machines offer SSH access and per-second billing with no minimum commitment. Stop and release the compute when the pilot ends, and verify deletion or retention separately for every stored copy. Obtain the required production approval before expanding the workload.
Assemble the artefacts, then request the DPA and documentation before you book the review.