Three layers to declare, not one

GPU capacity is usually the last resource in an otherwise declared estate that still gets clicked into existence. Specialist GPU clouds ship a REST API and a CLI long before they ship a Terraform provider, so the node with the highest hourly rate on the invoice is provisioned by hand, configured by hand, and attached to a cluster that everything else manages as code. That is exactly where drift concentrates.

The reason drift hurts more here than on a stateless web tier is that a GPU node is not one resource. It is three, and they fail independently:

  • Node: instance identifier, hardware profile, GPU count and SSH access
  • Host and container: compatible driver, kernel, CUDA image and NVIDIA Container Toolkit versions
  • Scheduler: GPU resource and device-plugin configuration where Kubernetes is used. If requests are specified alongside limits, GPU request and limit values must match

Collapse the three layers into one manual step and version matching becomes a per-node artisanal process. Replicas that were identical on day one stop being identical after the second rebuild, and the failure shows up later as a container that cannot see the GPU at all. NVIDIA built the GPU Operator precisely because configuring drivers, container runtimes and device plugins by hand across a fleet is difficult and prone to errors. The rest of this guide declares each layer in turn, then handles the case where the GPU cloud behind the node layer exposes only an API.

Pinning driver, CUDA and container runtime

The driver, kernel and container CUDA runtime must be compatible. Pin a tested combination and preserve the image or package sources needed to rebuild it. Pinning alone does not prove compatibility, so validate GPU visibility and a representative workload after changes.

The following shell pattern requires an explicit package version that your distribution repository provides. Set NVIDIA_CONTAINER_TOOLKIT_VERSION to the version you have tested before running it.

#!/usr/bin/env bash
set -euo pipefail
: "${NVIDIA_CONTAINER_TOOLKIT_VERSION:?Set a tested package version}"
sudo apt-get install -y \
  "nvidia-container-toolkit=${NVIDIA_CONTAINER_TOOLKIT_VERSION}" \
  "nvidia-container-toolkit-base=${NVIDIA_CONTAINER_TOOLKIT_VERSION}" \
  "libnvidia-container-tools=${NVIDIA_CONTAINER_TOOLKIT_VERSION}" \
  "libnvidia-container1=${NVIDIA_CONTAINER_TOOLKIT_VERSION}"

The same discipline applies to the driver. Install it from the distribution's package manager with a pinned package version rather than a.run installer, so a rebuild reproduces the same driver against the same kernel. If the nodes run Kubernetes, the GPU Operator collapses this layer into a Helm value: it automates the management of the NVIDIA driver, the container toolkit, the Kubernetes device plugin, node labelling and DCGM monitoring as one operator-managed stack, so pinning becomes a chart value in Git rather than a per-node script.

  • Host layer: pinned driver package and kernel pair, declared in the image or provisioning script.
  • Container layer: pinned NVIDIA Container Toolkit version, exactly as the install guide's own export shows.
  • Cluster layer: GPU Operator chart values (or the device plugin manifest) pinned in the same repository as the node layer.

Declaring nodes where a provider exists

A native provider defines which changes can update an instance and which require replacement. Read its schema and the actual plan. Do not assume every GPU service offers the same resizing or replacement behaviour.

Treat a hardware change as disruptive until the provider documents otherwise. Check job checkpointing, replacement capacity and the order of operations before applying it.

Check hardware availability before the apply, not after it fails. The VM API exposes GET /api/v2/external/vms/availability, which returns the hardware profiles you are allowed to provision. Treat the check as a precondition that filters out impossible plans, not a capacity guarantee: a listed profile is not a reservation, and capacity can be gone by the time the apply runs.

curl --fail-with-body --silent --show-error \
  https://api.lyceum.technology/api/v2/external/vms/availability \
  -H "Authorization: Bearer ${LYCEUM_API_KEY:?Set LYCEUM_API_KEY}"
  • Declare hardware arguments (profile, GPU count) as resource attributes so a change plans a replacement.
  • Read the plan for the word replace next to any GPU node before approving it.
  • Query the availability endpoint as a plan-time precondition, remembering that a listed profile is not a capacity reservation.

Recording an API-created VM without pretending it is managed

terraform_data stores input and replacement triggers in Terraform state. It does not discover a VM ID, refresh remote VM status or reconcile drift. Provisioners can invoke scripts, but the script author must implement those behaviours. Prefer a maintained native provider when one meets the service’s API contract.

Lyceum documents create, status, list and terminate operations. Creation returns vm_id; status polling recognises ready or running as successful and failed or error as failure. Keep that returned ID in durable job inventory immediately. If creation times out before the response, reconcile the remote inventory before retrying: the request may already have created billable capacity.

This runnable Terraform configuration records an already-created VM supplied by an external provisioning workflow. It creates no cloud resource and makes no HTTP calls. Run terraform init, then terraform plan with an actual recorded VM ID and hardware profile.

terraform {
  required_version = ">= 1.4.0"
}

variable "vm_id" {
  type        = string
  description = "Existing VM ID returned by the provisioning workflow"
  validation {
    condition     = length(trimspace(var.vm_id)) > 0
    error_message = "Provide the recorded VM ID."
  }
}

variable "hardware_profile" {
  type = string
}

resource "terraform_data" "gpu_inventory" {
  input = {
    vm_id            = var.vm_id
    hardware_profile = var.hardware_profile
  }
}

output "recorded_vm_id" {
  value = terraform_data.gpu_inventory.output.vm_id
}

The record is useful for downstream configuration, but changing or destroying it does not resize or terminate the VM. Keep create, readiness checks and termination in a documented external workflow until a provider owns the remote lifecycle. Store credentials outside Terraform input and state.

  • Before create, check availability and record a unique operation in a durable inventory
  • Persist the returned vm_id before polling readiness; use a deadline and handle failure states
  • On an uncertain create result, reconcile the provider inventory before another create request
  • On failure, use the recorded ID to inspect and explicitly terminate unwanted capacity
  • After termination, confirm the remote instance state and billing. Removing a Terraform record does not stop a VM

Remote state and locking for expensive resources

State handling matters more when a bad apply costs a GPU node. By default Terraform stores state locally in terraform.tfstate, which means every team member must hold the latest state and nobody else must run Terraform at the same time. Remote state writes the state data to a shared backend such as HCP Terraform, Amazon S3, Azure Blob Storage or Google Cloud Storage, so the whole team works from one state.

Locking comes with the backend. If the backend supports it, Terraform locks the state for all operations that could write state, which prevents concurrent applies from corrupting it; if acquiring the lock fails, Terraform does not continue. For resources where a duplicate apply means a second H100 node billing from the minute it enters running, that is not a nicety.

Use remote state with supported locking and review plans that affect real node resources. Separately protect the external VM inventory and serialise provisioning operations. Terraform state locking does not automatically lock an independent script or the provider’s API.

  • Store state remotely (S3, Azure Blob, GCS, HCP Terraform) so the team shares one state instead of local terraform.tfstate files.
  • Rely on automatic state locking on every write operation; do not disable it with -lock=false.
  • Gate any plan containing a replace on a GPU node behind human review.
  • Use one org-scoped lk_ key per organization in CI; the key pins the org, no extra header required.

Guarding a node that holds a running job

The worst Terraform failure on a GPU estate is not drift, it is an accidental destroy of a node holding a multi-day training job. Terraform's lifecycle rules exist for exactly this, and each is documented behaviour, not opinion:

  • prevent_destroy rejects destruction while its configuration remains present; removing the resource block removes that protection
  • create_before_destroy may require overlapping paid capacity. It also prevents destroy-time provisioners from running on that resource
  • Destroy-time provisioners can also be skipped for tainted resources or removed configuration. Do not depend on them as the only VM cleanup path
  • Use ignore_changes only for fields another documented owner manages; it can hide meaningful drift

For a native resource, choose lifecycle settings according to the provider and workload. For the inventory-only terraform_data example, those settings protect or replace the record, not the real VM. HashiCorp documents the provisioner lifecycle caveats explicitly.

Verifying a rebuild reproduces the same node

Validate rebuilds on disposable test capacity after material image or driver changes. Budget the test, preserve results, and confirm termination afterwards. Do not destroy a production training node to test reproducibility.

  • nvidia-smi shows the expected GPU count and model on the fresh node.
  • The pinned driver version and NVIDIA Container Toolkit version match the previous node exactly.
  • A test container sees the GPU, confirming the toolkit wiring survived the rebuild.
  • Results are pulled off first: local disk on a VM is wiped on termination, so anything you want to keep is scp'd off, pushed to Git, or written to storage before terminating.

If the rebuild produces a node with a different driver or toolkit version, the pinning in the previous sections has a gap, and it is better to find that in a test than during a production replacement. This is also the point to confirm what stays outside Terraform: model weights, datasets and anything whose lifecycle is not the node's belong in object storage and your training pipeline, not in the node's configuration. Terraform declares the node; the node is disposable by design.

Lyceum GPU VMs expose a documented REST API and CLI. Confirm available hardware, topology and price before creating an instance. Until remote lifecycle management is implemented and tested, keep the API workflow and Terraform inventory boundary explicit.