The Fine-Tuning Fear: Are You Deploying or Providing?

You are ready to launch a fine-tuning run on an open-weight model, but legal puts the job on hold. The fear is straightforward: does modifying an open-source base model turn your engineering team into a General-Purpose AI (GPAI) provider under the EU AI Act, Regulation (EU) 2024/1689? If it does, your team inherits the provider obligations: detailed technical documentation, a copyright policy, and a published summary of the content used for training.

European lawmakers anticipated this friction. Recital 109 of the EU AI Act states that in the case of a modification or fine-tuning of a model, the obligations for providers of general-purpose AI models should be limited to that modification or fine-tuning, for example by complementing the already existing technical documentation with information on the modifications, including new training data sources. Separately, where a model meets the compute threshold, or it becomes known that it will be met, the provider must notify the Commission without delay and within two weeks at the latest under Article 52(1). To prevent routine engineering work from triggering regulatory lock-in, the Commission structured GPAI classification around quantifiable thresholds rather than subjective legal interpretations, and the resulting foundation model obligations are documented separately.

  • Deployer status: You adapt an existing GPAI model through prompt engineering, retrieval pipelines, or lightweight parameter-efficient tuning without crossing compute boundaries.
  • Downstream provider status: You execute substantial post-training modifications that fundamentally alter model capabilities or exceed the European Commission's compute thresholds.
  • Systemic risk tier: Cumulative training compute exceeds 10^25 FLOP, requiring mandatory notification under Article 52(1).

You do not need to guess your legal status or debate ambiguity in contract clauses. Under the Commission's Guidelines on the scope of the obligations for providers of general-purpose AI models, actors modifying or fine-tuning a model become providers only in exceptional circumstances, specifically when the modification or fine-tuning uses more than one-third of the original model's training compute. For model-training teams, the question is not a matter of legal opinion; it is an arithmetic calculation based on parameters, dataset tokens, and floating-point operations.

The Baseline: What Makes a GPAI Model?

The AI Act defines a general-purpose AI (GPAI) model by its generality and its capability to competently perform a wide range of distinct tasks. To translate that qualitative definition into an objective test for engineering teams, the Commission's guidelines anchor an indicative classification criterion to cumulative training compute measured in floating-point operations (FLOP): a model qualifies where training compute exceeds 10^23 FLOP and it can generate language (text or audio), text-to-image, or text-to-video.

Classification TierCompute ThresholdLegal BasisRegulatory Scope
Indicative GPAI Baseline> 10^23 FLOPCommission Guidelines on the scope of GPAI obligationsModels that can generate language (text or audio), text-to-image or text-to-video, subject to the Article 53 documentation, copyright policy and training-data summary duties
Systemic Risk GPAI> 10^25 FLOPRegulation (EU) 2024/1689, Art. 51(2)Models presumed to have high-impact capabilities, whose providers must additionally carry out model evaluation, systemic risk assessment and mitigation, incident reporting and cybersecurity protection

The Commission's guidelines treat a model as a general-purpose AI model where its training compute exceeds 10^23 FLOP and it can generate language (text or audio), text-to-image, or text-to-video. This is an indicative criterion rather than an absolute rule: models above it may exceptionally not qualify if they lack significant generality, and models below it may still qualify if they can competently perform a wide range of tasks. Crossing the line triggers the GPAI model obligations in Article 53, requiring providers to maintain technical documentation, implement a copyright policy, and publish a training data summary.

By contrast, the 10^25 FLOPs threshold defined in Article 51(2) is reserved for systemic-risk frontier models, bringing stricter requirements under Article 55 such as continuous adversarial testing and model evaluation. For machine learning teams planning downstream fine-tuning runs, establishing these baseline compute boundaries is the necessary starting point before measuring whether a subsequent modification shifts regulatory status.

The One-Third Rule for Downstream Modifiers

Modifying an open-weight foundation model does not automatically transfer the regulatory obligations of a primary model provider to your engineering team. The Commission's position is that not every modification or fine-tuning of a general-purpose AI model should be treated as creating a new model: actors modifying or fine-tuning a model become providers only in exceptional circumstances, specifically when the modification or fine-tuning uses more than one-third of the original model's training compute.

To establish where routine fine-tuning ends and provider status begins, the European Commission introduced a quantitative compute metric in its Guidelines on the scope of the obligations for providers of general-purpose AI models. At paragraph 63 the guidelines state that "an indicative criterion for when a downstream modifier is considered to be the provider of a general-purpose AI model is that the training compute used for the modification is greater than a third of the training compute of the original model". The same criterion is cross-referenced in the guidelines' notification rules as "the threshold laid down in paragraph 60". Read it together with paragraph 62: the modifier becomes the provider only if the modification leads to a significant change in the model's generality, capabilities or systemic risk. When the base model's pre-training compute is publicly documented, applying that threshold is a direct ratio.

  • Known original training compute: the indicative criterion is that "the training compute used for the modification is greater than a third of the training compute of the original model" (paragraph 63).
  • Unknown base GPAI compute: where the modifier cannot be expected to know the original figure and cannot estimate it, paragraph 64 replaces the threshold with a third of the threshold for a model being presumed to be a general-purpose AI model, currently 10^23 FLOP.
  • Unknown systemic model compute: where the original model is a general-purpose AI model with systemic risk, paragraph 64 substitutes a third of the threshold for a model being presumed to have high-impact capabilities, currently 10^25 FLOP under Article 51(2).

Paragraph 64 of the guidelines resolves scenarios where the upstream developer does not disclose pre-training metrics. Where the downstream modifier cannot be expected to know that value and cannot estimate it, the threshold is replaced with a third of the threshold for a model being presumed to be a general-purpose AI model (currently 10^23 FLOP), or, if the original model is a general-purpose AI model with systemic risk, a third of the 10^25 FLOP threshold in Article 51(2). If your run stays below these computational boundaries, you avoid inheriting the full GPAI foundation model obligations, remaining classified as a downstream deployer rather than an upstream provider.

How to Measure Your Fine-Tuning Compute

To determine whether a post-training modification crosses the regulatory threshold, you must turn the legal question into an arithmetic calculation of total floating-point operations (FLOP). The Commission's guidelines point providers to "the Annexes A.1 and A.2 to these guidelines for how to estimate training compute", and say that providers "should estimate the cumulative amount of training compute that they will use before starting" the large pre-training run. Regulators do not expect cycle-level telemetry from every individual CUDA kernel: a notification must describe "the approach used to estimate this amount of compute, including approaches used to make approximations where precise information is not available".

Hardware-Based vs Architecture-Based Approaches

  • Hardware-based estimation: Calculates compute directly from hardware allocation using the formula: Total FLOP = (Number of GPUs) × (Runtime in seconds) × (Datasheet dense peak FLOP/s per GPU at execution precision) × (Model FLOPs Utilisation, or MFU).
  • Architecture-based estimation: Calculates compute from model dimensions and dataset size using the standard forward-backward training multiplier for dense transformers, C ≈ 6ND, where N is the parameter count and D the number of training tokens.

For machine learning engineering teams deploying jobs on dedicated infrastructure, the hardware-based approach is almost always the simplest to calculate. It relies on three straightforward operational variables that you already monitor: the number of provisioned accelerators, the elapsed runtime of the training container, and the vendor's published peak throughput at your target precision (such as dense FP8 or BF16). When you provision compute through GPU training infrastructure, your invoice records the GPU model and exact GPU-hours consumed. Multiplying those logged operational quantities by your workload's measured MFU yields a defensible, audit-ready FLOP total well within the 30% margin.

The Math: Standard Fine-Tuning vs the Threshold

Determining whether a fine-tuning job reclassifies your team from a downstream deployer to a general-purpose AI (GPAI) provider under the EU AI Act does not require subjective legal guesswork; it is a straightforward arithmetic problem. Article 51(2) of the AI Act presumes a model to have high-impact capabilities when the cumulative amount of computation used for its training measured in floating point operations is greater than 10^25. For downstream modifications, the relevant comparison is one-third of the base model's training compute, measured against the compute of your modification alone.

Calculating FLOPs for an 8-GPU Node

To establish your compute footprint, calculate total operations using hardware throughput: GPU count × duration in seconds × dense peak FLOP/s per GPU at your execution precision × Model FLOPs Utilisation (MFU). Consider a typical workload: fine-tuning an open model across an 8-GPU node running continuously for seven full days. NVIDIA's H100 SXM datasheet gives a BF16 Tensor Core rate of 1,979 TFLOPS with sparsity, which is twice the dense rate of 989 TFLOPS; a training-compute estimate must use the dense figure, because using the with-sparsity number halves the apparent GPU-hours and understates the estimate twofold.

Hardware setup (dense BF16 peak rate)Duration of continuous trainingTheoretical peak compute, order of magnitude
8× NVIDIA H100 SXM at 989 TFLOPS dense BF16168 hours (one week)Of the order of 10^21 FLOP, using GPU count × runtime × dense peak FLOP/s × utilisation
32× NVIDIA H100 SXM at 989 TFLOPS dense BF16336 hours (two weeks)Of the order of 10^22 FLOP on the same hardware-based method

Even assuming perfect hardware efficiency, a week-long run on eight H100s lands at an order of magnitude around 10^21 FLOP under the hardware-based method of GPU count × runtime × dense peak FLOP/s × utilisation, which is well below a third of the 10^23 FLOP indicative criterion and orders of magnitude below the 10^25 FLOP systemic-risk presumption. Because real Model FLOPs Utilisation is a fraction of peak rather than all of it, the actual figure for your job is lower still. For teams undertaking standard enterprise runs when fine-tuning a 70B model, the arithmetic shows a wide margin between the workload and either threshold. Only a sustained multi-week job on tens of GPUs starts to approach the modification line.

Billing and Infrastructure: The Cloud Provider's Role

When preparing a fine-tuning run, engineering teams often wonder what compliance telemetry their cloud infrastructure provider generates. The short answer is that an infrastructure provider only sees and meters raw resource allocation. We provision compute nodes, manage network interconnects, and execute the containers you deploy. A cloud host does not inspect loss curves, audit token distributions, or determine whether your architectural modification legally creates a new general-purpose AI model.

The connection between your cloud bill and regulatory compliance is purely arithmetic. Products like On-demand GPU VM bill per second and training products are priced per GPU-hour. The GPU-hours consumed and the GPU model a job ran on are the exact quantities the customer is billed on. Under standard hardware-based compute estimation methodologies, GPU-hours multiplied by nominal GPU processing throughput represents the starting point for calculating your total operations.

  • Infrastructure provisioning: Supplying raw GPU hardware, host environments, and container runtimes such as Serverless Training.
  • Billing metrics: Invoicing strictly on consumed GPU-hours, instance duration, and specific hardware SKUs.
  • Engineering compliance: Logging precise runtime parameters, measuring actual Model FLOPs Utilisation (MFU), and determining your regulatory status under the EU AI Act.

Lyceum sells the compute and provisions raw capacity; it does not report, calculate, certify, or supply floating-point operation totals or AI Act legal classifications. Your team owns the model architecture, tracks the actual training efficiency, and calculates your cumulative compute. We focus on providing high-performance infrastructure, and you calculate your own numbers to settle your regulatory standing.

Execute Your Run with Serverless Training

Once the arithmetic confirms your fine-tuning run remains orders of magnitude beneath the general-purpose AI compute thresholds established in Regulation (EU) 2024/1689, the focus shifts from regulatory classification to workload execution. Lyceum sells the compute capacity that fine-tuning runs consume, not legal determinations. We do not calculate floating-point operations, certify compliance profiles, or assign provider roles under the EU AI Act. Instead, we deliver the low-latency hardware execution that model-training teams require. For teams modifying open-weight architectures, Serverless Training provides direct access to dedicated compute without cluster management overhead.

Containerised Deployment Across European Facilities

Serverless Training eliminates cluster maintenance overhead and idle resource costs. You package your PyTorch or Hugging Face training script into a Docker container, define your hardware requirements, and submit the job. The orchestration layer pulls your container image directly from your repository and initializes the runtime on dedicated hardware across European data centres in Paris and Finland. Because the system bills purely on GPU-hours for the exact accelerator model provisioned, you retain precise telemetry for both cost accounting and your internal compute estimation logs.

  • Direct registry integration: Pull container images directly from Amazon Elastic Container Registry (ECR), Google Artifact Registry (GAR), or Docker Hub.
  • Fast runtime initialization: Automated containerisation and node preparation get your training run started in under 60 seconds.
  • S3-compatible persistent storage: Mount standard object storage buckets directly into your execution environment to stream training datasets and write model checkpoints.
  • Architecture selection: Deploy across modern enterprise accelerators including NVIDIA L40S, A100, H100, H200, B200, and B300 hardware.

Stop letting misplaced compliance anxiety stall your engineering roadmap. When your modification compute is a fraction of the regulatory boundary, execution speed is what matters. Package your container, define your GPU target, and submit your job to Serverless Training today.