AI This article was created with the help of AI.
Bring-Your-Own-Weights: Wan 2.2 on Cloud GPUs
If you are searching for a managed serverless API endpoint to call Wan 2.2, the direct answer is that no such pre-hosted text-to-video endpoint exists on Lyceum. Our serverless catalogue provides Wan Image for text-to-image workloads and Video Replace for subject and face replacement in existing footage, but text-to-video generation remains a bring-your-own-weights workload. Running models like Wan 2.2 requires provisioning dedicated GPU infrastructure, pulling the weights, and managing the runtime environment directly.
The Wan model series released its architecture and weights under an open license, establishing a competitive open-weight standard for video generation across 1.3B and 14B parameter variants. Because the weights are public, engineering teams are not locked into proprietary, rate-limited commercial video APIs. However, bringing these models to cloud GPUs fundamentally shifts the economic and operational paradigm from output-metered API pricing to time-based infrastructure billing.
When deploying open-weight video models on On-demand GPU VM instances, you pay strictly for the GPU seconds your instances consume rather than a flat fee per generated second. Calculating your true cost per clip therefore requires understanding the memory profile of the diffusion pipeline, matching the workload to the correct GPU hardware tier, and ensuring that compute resources do not sit idle between rendering batches.
- Managed API Model: You send a prompt over HTTP and pay a fixed price per output clip, with zero infrastructure management but no control over execution precision, batch scheduling, or pipeline caching.
- Bring-Your-Own-Weights Model: You provision raw GPU compute over SSH or run private containers, pull model checkpoints directly, tune VRAM offloading strategies, and pay for active execution time.
Video Diffusion Memory vs. LLM Memory
Engineers accustomed to serving large language models often attempt to apply autoregressive memory sizing rules to video diffusion pipelines. In LLM inference, memory is dominated by static weight allocations and dynamic Key-Value (KV) caches that scale linearly with context length and batch size. Video diffusion architectures, such as the 3D Diffusion Transformers (DiTs) powering Wan 2.2, operate on entirely different mathematical dynamics.
In video diffusion, memory requirements scale with spatial resolution (height and width) and temporal sequence length (total frame count). The denoising network operates on 3D video tokens that carry both spatial and temporal information, and the VAE pass that decodes those latents back to pixels is a separate allocation on top of the denoise loop, which is why frameworks expose sliced and tiled decoding specifically to cut peak activation memory at that stage. Which of the two stages sets your actual peak is something you have to measure rather than assume. Doubling the duration or increasing output from 480p to 720p therefore expands intermediate activation memory sharply during the forward passes of the denoising backbone.
The Multi-Stage Video Pipeline Allocation
A standard video diffusion pipeline divides memory consumption across four distinct components during execution:
A critical operational detail is that the number of denoising steps does not expand your VRAM footprint. Running 30 steps versus 50 steps consumes identical peak memory; it simply extends the wall-clock execution time on the GPU. Frame count and resolution dictate whether a clip fits in VRAM without triggering a CUDA Out-of-Memory (OOM) crash, while the step count dictates what the clip costs in compute time.
| Dimension | LLM Inference | Video Diffusion (Wan 2.2) |
|---|---|---|
| Primary Dynamic Consumer | KV Cache scaling with context | Spatiotemporal attention activations and VAE decode |
| Scaling Axes | Batch size and token count | Frame count, height, and width |
| Peak Memory Moment | Maximum context fill during generation | Intermediate denoising steps or VAE pixel decoding, whichever profiles higher |
| Step / Token Impact | Tokens increase memory linearly | Steps increase execution time, not peak VRAM |
Sizing Procedure for Wan 2.2 Workloads
Rather than relying on static VRAM sizing tables that quickly become outdated as libraries update, engineering teams should implement an empirical profiling routine to size video diffusion workloads accurately. Video generation memory overhead is heavily sensitive to framework-level optimizations such as FlashAttention, sequence chunking, and VAE tiling.
A Step-by-Step Memory Profiling Routine
To size a Wan 2.2 deployment without risking runtime crashes in production, follow this profiling sequence:
Benchmarking indicates that running full unquantized FP16 pipelines for 14B video models at 720p resolution without offloading requires between 65 GB and 80 GB of dedicated VRAM. Applying FP8 transformer quantization and offloading the text encoder brings the operational threshold down to 16 GB to 24 GB, making hardware selection a direct trade-off between engineering complexity and raw card capacity.
Matching Workloads to the GPU Fleet
Once your pipeline's peak memory requirements are profiled, the next step is selecting the appropriate hardware tier. An enterprise GPU fleet typically spans a mid-capacity workstation-class card (NVIDIA L40S), two data centre accelerators that both publish 80 GB of HBM (NVIDIA A100 and NVIDIA H100), and higher-capacity Hopper and Blackwell parts for jobs whose peak allocation exceeds what an 80 GB card can hold.
The 48 GB tier (NVIDIA L40S) is well-suited for quantized Wan 2.2 workloads (such as FP8 or GGUF) and 480p generation pipelines utilizing CPU text encoder offloading. However, because the L40S relies on GDDR6 memory rather than high-bandwidth HBM, per-step iteration times are longer than on data centre accelerator architectures.
Evaluating the 80 GB Tier: A100 vs. H100
Both the NVIDIA A100 and NVIDIA H100 provide 80 GB of VRAM, meaning that from a pure memory-fit perspective, any video diffusion job that fits on an H100 will also fit on an A100. The distinction between them is memory bandwidth and compute throughput rather than capacity.
The A100 80 GB SXM delivers 2,039 GB/s of HBM2e bandwidth, whereas the H100 SXM delivers 3.35 TB/s of HBM3 bandwidth alongside fourth-generation Tensor Cores and native FP8 Transformer Engine support. Because diffusion loops are memory-bandwidth-bound during iterative latent passes, the H100 generates identical clips substantially faster, reducing the total seconds the GPU must be held per batch. Sizing considerations mirror those in our guide on VRAM requirements for large architectures.
| GPU Model | VRAM Capacity | Memory Type | Memory Bandwidth | Workload Profile |
|---|---|---|---|---|
| NVIDIA L40S | 48 GB | GDDR6 | 864 GB/s | Quantized FP8/GGUF, 480p batches, CPU offload |
| NVIDIA A100 | 80 GB | HBM2e | 2,039 GB/s | Unquantized 720p baseline, cost-effective rendering |
| NVIDIA H100 SXM | 80 GB | HBM3 | 3,350 GB/s | High-throughput 720p production, native FP8 support |
| NVIDIA H200 | 141 GB | HBM3e | 4,800 GB/s | Unquantized 720p/1080p, long-sequence multi-frame clips |
| NVIDIA B200 | 192 GB (180 GB usable) | HBM3e | 8,000 GB/s | Maximum throughput parallel video generation |
For extended video sequences or unquantized 1080p rendering where peak memory exceeds 80 GB during temporal attention or VAE decoding, the NVIDIA H200 offers 141 GB of HBM3e memory at 4.8 TB/s, nearly double the capacity of the H100, which removes out-of-memory risk without requiring complex pipeline sharding. When targeting the NVIDIA B200, note that while 192 GB is the theoretical HBM3e capacity, production teams should budget memory allocations against the 180 GB capacity shipped in standard HGX configurations. Choosing the right tier follows our GPU cloud provider evaluation criteria.
Multi-GPU Interconnects and Scaling Constraints
When a video generation workload exceeds the physical memory capacity of a single GPU, or when batch throughput demands distributed processing, engineering teams must decide between pipeline sharding and parallel worker distribution. Sharding a single diffusion pipeline across multiple GPUs introduces communication overhead that depends directly on the underlying hardware interconnect.
On-demand GPU VM instances allow provisioning from 1 to 8 GPUs per virtual machine. On SXM and HGX platforms equipped with NVIDIA NVLink, inter-GPU communication operates at up to 900 GB/s bidirectional bandwidth, enabling tensor-parallel or context-parallel diffusion execution with minimal latency penalties.
However, interconnect topology varies significantly across GPU tiers. The NVIDIA L40S 48 GB tier does not feature NVLink support. Multi-GPU L40S configurations communicate strictly over PCIe buses. Attempting to shard a single latent tensor or split transformer layers across PCIe creates severe communication bottlenecks that degrade generation speed.
While lightweight models like the Wan 1.3B baseline require only 8.19 GB VRAM and run comfortably on single consumer-grade cards, scaling 14B and Mixture-of-Experts architectures across multi-GPU nodes requires strict attention to interconnect bandwidth. For most production pipelines, running independent single-GPU generation workers across an 8-GPU node is far more efficient than splitting a single clip across multiple cards.
The Cost-per-Clip Calculation
On rented GPU infrastructure, the cost to produce an AI video clip is an identity governed by wall-clock render time, cluster throughput, and infrastructure billing increments rather than a static per-output fee. The fundamental cost equation is:
Cost per Clip = ((Wall-Clock GPU Seconds Held / Clips Produced in Window) * Per-Second GPU Rate) + Storage Egress Overhead
Because On-demand GPU VM instances are billed with per-second billing and no egress fees, you do not pay for rounded-up fractional hours or outbound bandwidth when retrieving rendered video assets. Live per-second GPU rates and storage rates can be verified directly at lyceum.technology/pricing.
Worked Example: 720p Clip Rendering
To illustrate the arithmetic, consider generating a 5-second 720p clip using Wan 2.2 on an NVIDIA H100 SXM instance. Community benchmarks record average generation times of approximately 10 to 12 minutes (600 to 720 seconds) for a standard 50-step diffusion schedule on a single card.
| Pipeline Configuration | Target Output | Avg Render Time (H100) | GPU Seconds Consumed | Cost Formula |
|---|---|---|---|---|
| Wan 2.2 (FP8, 30 steps) | 480p (5 seconds) | ~4 minutes | 240s | 240 * (Per-Second GPU Rate) |
| Wan 2.2 (FP16, 50 steps) | 720p (5 seconds) | 10 to 12 minutes | 600 to 720s | 600-720 * (Per-Second GPU Rate) |
| Wan 2.2 (FP16, 50 steps) | 1080p (5 seconds, H200) | Longer than 720p; profile your own pipeline | Measured per job | Measured seconds * (Per-Second H200 Rate) |
Because video files are substantially larger than text embeddings, storing and transferring high-resolution clips on standard hyperscalers often incurs hidden egress surcharges. Free egress and S3-compatible object storage ensure that the storage term in the cost equation remains strictly limited to static disk retention without data transfer penalties.
Utilization and Scale-to-Zero
In production video generation, the variable that dominates your total infrastructure spend is not the hourly hardware rate: it is hardware utilization. Video generation workloads are inherently bursty. Editorial teams, creative tools, and automated batch rendering queues submit workloads in sporadic spikes, followed by hours of zero activity.
If a GPU instance sits idle overnight or across weekends waiting for jobs, your effective cost per clip multiplies. The billed hours stay the same whether the card is rendering or waiting, so a cluster held up between bursts pays the full rate for every idle second, and no comparison of headline hourly rates recovers that money.
By aligning memory profiling, hardware tier selection, and automated scale-to-zero infrastructure across European data centres in Paris and Finland, engineering teams can operate open-weight video diffusion pipelines at predictable, optimized unit economics.