One provider is a single point of failure

When inference sits in the user request path, a provider outage can interrupt the features that depend on it. Cached answers, queues or reduced functionality may keep other parts of the product available. Design the fallback around the actual failure your users would experience.

A second endpoint can reduce exposure to provider outages, but it adds cost and operational work. Check whether the providers share an upstream service, region or network dependency before treating them as independent.

Architecture PatternOutage ImpactFailover LatencyOperational Overhead
Single providerDependent features may fail or waitDepends on recovery or manual interventionOne provider to operate
Dual-provider active-passiveBackup can accept eligible requestsFailure detection plus backup response timeTest backup capacity and recovery
Dual-provider active-activeHealthy capacity can receive new trafficDepends on detection, routing and spare capacityMonitor both providers and limit load

Building resilience into your inference pipeline does not require rewriting your application layer. When structured correctly, implementing an active-passive or active-active failover architecture between two independent hosts protects your product from upstream infrastructure disruptions.

When the same model exists on both sides

The simplest case is the same published checkpoint available from two independent providers. Compare model cards, checkpoint revisions and licences. Open weights alone do not establish compliance with the Open Source AI Definition or give access to a provider’s serving runtime.

The same checkpoint reduces one source of variation, but does not guarantee equivalent outputs. Quantisation, chat templates, tokenizers, sampling defaults and serving engines can differ. Run your application’s evaluations and validate schemas and tool calls on both paths.

Some proprietary models are available through multiple platforms, although these routes may share an upstream dependency. Confirm the exact version and failure domain. If your backup uses another model family, evaluate that as a behaviour change as well as an infrastructure change.

Putting two endpoints behind one client

For compatible endpoints, keep provider-specific base URLs, keys and model IDs in configuration. Route requests through one wrapper and normalize only the fields your application needs. Compatibility still needs testing on each provider.

Open-source serving stacks behave differently depending on how they are configured. The PagedAttention paper describes an attention algorithm inspired by classical virtual memory and paging techniques in operating systems, on top of which the authors built vLLM, a serving system achieving near-zero waste in KV cache memory and flexible sharing of that cache within and across requests. The project itself lists an OpenAI-compatible API server alongside continuous batching, prefix caching and structured-output generation as separate features of the engine, and its engine arguments control the behaviour of the engine for online serving. Different providers set those runtime defaults differently. Before deploying automated failover into production, you must verify and document four operational divergence areas:

  • Sampling parameter parsing: Verify whether both endpoints strictly adhere to temperature, top_p, top_k, and repetition_penalty, or silently ignore unsupported arguments.
  • Structured output handling: Test whether JSON schema constraints and tool calling operate identically via open models with reliable function calling across both hosts.
  • Streaming: Check chunk shapes, completion markers and failures after the first token; never splice a second generation into a partial answer.
  • Usage: Check whether streaming requires stream_options.include_usage, and handle missing usage after interrupted requests.

Use the same representative prompts, schema checks and tool scenarios against both endpoints. Also test unavailable backups and shared upstream failures.

Designing the trigger for each failure type

A resilient failover implementation requires explicit classification of upstream failure modes. Treating all HTTP errors identically leads to retry storms, exacerbated rate limits, and cascade failures across both providers. Your client wrapper must inspect response codes and connection states before deciding whether to retry locally, switch providers, or fail immediately.

Distinguishing HTTP status codes

HTTP 429 indicates rate limiting; inspect the provider’s error body for its cause. Honour Retry-After when retrying that endpoint, and use bounded backoff with jitter. Fail over only when the backup has capacity, the request is eligible and the overall deadline allows it.

Treat selected 500, 502, 503 and 504 responses as possible transient failures. A 504 is a gateway timeout; a client timeout can occur without any HTTP response. Do not blindly replay 400 or 422 validation errors, or 401 and 403 authentication failures. Provider-specific errors need explicit handling rather than assuming all endpoints fail identically.

Give each provider a circuit breaker and a concurrency limit. After a defined failure threshold, stop new requests to that provider temporarily. Probe recovery with limited traffic. Coordinate application retries with SDK retries so they do not multiply attempts or exceed the end-to-end deadline.

Avoiding duplicate work on a timeout

Handling client-side timeouts is the most complex failure mode in distributed inference routing. When a client aborts an HTTP connection after reaching its Time-to-First-Token (TTFT) deadline, the upstream GPU cluster may still be processing the prefill phase or generating tokens in its continuous batching queue.

A disconnected client does not prove that generation stopped upstream. A backup request can overlap with the original and both may be billed. Prompt length, load, cache state and the provider’s cancellation policy affect the result; a benchmark cannot establish cancellation or billing behaviour for your application.

  1. Bound output with the provider-supported token limit, allowing room for reasoning where needed.
  2. Estimate prompt size with the matching tokenizer and measure latency by prompt length and load.
  3. Cancel using the client’s supported API, then verify whether the provider stops generation and billing; closing a connection alone is not proof.
  4. Set separate first-token, stream-idle and overall deadlines from measured latency and your user-facing budget.

These controls reduce duplicate work; they cannot guarantee one bill across independent providers. Use an application request ID to deduplicate accepted results and side-effecting tool actions. Provider idempotency keys, where supported, do not automatically deduplicate across providers. After streaming starts, report an interruption or explicitly restart the answer rather than silently replaying it.

Failover can move processing across a border

Failover can change where prompts, outputs and logs are processed. Define the allowed countries and approved processors before enabling a backup. A European company address or API gateway location does not establish where inference runs.

The General Data Protection Regulation (GDPR) has additional conditions for transfers of personal data outside the European Economic Area (EEA). Such transfers are not automatically unlawful: an applicable adequacy decision or appropriate safeguards may support them. EU-only contractual commitments can be stricter. Confirm the processing agreement, subprocessors and transfer arrangement for both routes.

Start with the model catalogue, then obtain confirmation of actual serving and fallback regions. Use an allowlist matching your contractual requirements. If no approved backup is available, fail closed or queue the request. Matching provider region labels alone does not establish compliance.

Forcing a failure to test the path

An untested failover mechanism is merely an untested assumption. Waiting for an unscheduled upstream outage to validate your multi-provider routing logic frequently uncovers unhandled exceptions, tokenizer discrepancies, or missing environment variables under production load.

  • Simulate 429 and 503 responses in staging, including Retry-After and an unavailable backup.
  • Inject delays before and after the first token; verify deadlines, partial-answer handling and bounded attempts.
  • Use synthetic invalid credentials to verify that 401 errors alert operators rather than causing uncontrolled retries.
  • Replay synthetic or approved evaluation prompts on both providers and compare quality, latency and schemas.
  • Test circuit recovery, backup capacity and tool-action deduplication under concurrent requests.

Lyceum provides an OpenAI-compatible serverless inference API for open-weight models with per-token billing. Confirm serving and fallback regions before including it in a residency-restricted route. Check current service status and agree any contractual availability requirements with sales. A second endpoint improves resilience only if its capacity and upstream dependencies have been checked.