Designing a fallback chain for model or region outages
For developers building against LLM or inference APIs, this guide shows how to design a fallback chain that survives model-specific failures and regional outages without turning your system into a retry storm. You’ll get concrete routing patterns, health checks, error handling rules, and deployable config examples for HTTP services and worker-based inference clients.
TL;DR — A fallback chain is a policy-driven router that tries a preferred model and region first, then degrades in a controlled order when it gets specific failure signals such as 429, 5xx, timeouts, or regional DNS/TLS/connect failures. The single most useful fix is to stop doing blind retries in the client and instead centralize failover rules with per-model health, idempotency, and a hard budget for latency and duplicate work. Reading time: ~7 min
What it is and where it sits
A fallback chain is the decision layer between your application and one or more model-serving endpoints. It answers: if the primary model or region is slow, rate-limited, or down, what do we try next, under what conditions, and when do we stop?
In a typical stack, it sits in one of three places:
- inside your application client library
- in a thin gateway service in front of model providers
- at the edge or service mesh layer for region-level routing, with application logic still deciding model-level fallback
What it replaces: ad hoc try/catch blocks, infinite retries, and “just point DNS somewhere else” incident responses.
What talks to it:
- your web app / API workers / background jobs
- observability systems feeding health state
- sometimes a cache layer for idempotent or replay-safe requests
What it talks to:
- primary inference endpoint in region A
- secondary endpoint in region B
- alternate model endpoint when the preferred model is unavailable or overloaded
A useful mental model: this is not just load balancing. Load balancing spreads normal traffic across healthy backends. A fallback chain encodes preference order and degradation policy when health or capability changes.
Client request
|
v
App/API server
|
v
Fallback router
|-- try #1: model=large region=us-east-1
| |-- success -> return
| `-- fail (429/5xx/timeout/connect error)
|
|-- try #2: same model region=eu-west-1
| |-- success -> return
| `-- fail
|
`-- try #3: smaller model region=eu-west-1
|-- success -> return degraded-response header
`-- fail -> return controlled error
The architecture context matters because model fallback and region fallback are different failure domains:
- model fallback handles capacity, model deployment bugs, quota exhaustion, or capability mismatch
- region fallback handles DNS, TLS handshake, network path, zonal/regional outages, and provider control-plane incidents
Do not collapse them into one rule set unless you want hard-to-debug behavior.
How it actually works
Mechanically, a fallback chain is four things:
- a ranked list of candidates
- failure classification rules
- a retry/failover budget
- state, usually a short-lived circuit breaker per model+region
One realistic end-to-end example
Assume your app normally sends chat-completion requests to model=reasoning-large in us-east-1. If that region is unavailable, you want to try the same model in eu-west-1. If the model itself is overloaded everywhere, you want to degrade to model=reasoning-medium in eu-west-1. Total extra latency budget: 2.5 seconds. No duplicate side effects allowed.
Step 1: the app sends a request with an idempotency key and deadline.
POST /v1/infer HTTP/1.1
X-Request-Id: 9d7f6b8a-3c7d-4f8e-a9b8-2fd8f7c1d0c1
Idempotency-Key: infer-01JV6YQ4Q6M8S8R4Y7YV7A8M2N
X-Deadline-Ms: 2500
Content-Type: application/json
The router stores attempt state keyed by request ID or idempotency key. This is what prevents replaying the same expensive request five times during a brownout.
Step 2: candidate #1 is selected.
[
{"model":"reasoning-large","region":"us-east-1","timeout_ms":1200},
{"model":"reasoning-large","region":"eu-west-1","timeout_ms":800},
{"model":"reasoning-medium","region":"eu-west-1","timeout_ms":500}
]
Step 3: the first upstream call fails with a connect timeout. That is a region/path failure, not a model-capability failure. The router records a failure against reasoning-large@us-east-1, opens a short circuit for, say, 30 seconds after N failures, and moves to candidate #2.
Typical client-visible logs:
{"ts":"2026-10-01T12:00:01Z","request_id":"9d7f6b8a-3c7d-4f8e-a9b8-2fd8f7c1d0c1","attempt":1,"model":"reasoning-large","region":"us-east-1","event":"upstream_error","error":"connect timeout","elapsed_ms":1203}
{"ts":"2026-10-01T12:00:01Z","request_id":"9d7f6b8a-3c7d-4f8e-a9b8-2fd8f7c1d0c1","attempt":2,"model":"reasoning-large","region":"eu-west-1","event":"dispatch"}
Step 4: candidate #2 returns HTTP 429 with Retry-After: 10. For an interactive request with a 2.5s deadline, you do not sleep 10 seconds. You classify this as capacity exhaustion and move to candidate #3 if policy allows degradation.
Example response shape:
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 10
{"error":{"type":"rate_limit","message":"capacity exceeded for model reasoning-large"}}
Step 5: candidate #3 succeeds. The router returns the response plus metadata so downstream systems know this was degraded.
HTTP/1.1 200 OK
X-Fallback-Attempts: 3
X-Served-Model: reasoning-medium
X-Served-Region: eu-west-1
X-Degraded: true
Content-Type: application/json
Step 6: metrics and alerting fire on the fallback, not just the final success. If you only monitor 200 rates, you will miss the outage until users complain about lower-quality answers.
The key concrete rules:
- fail over on connect timeout, TLS handshake failure, DNS resolution failure, HTTP 502/503/504, and provider 429 if your deadline cannot honor
Retry-After - do not fail over on 400/401/403/404; those are almost always caller/config errors
- be careful with 500: if the provider uses 500 for malformed prompts or unsupported options, inspect the body before retrying elsewhere
- preserve a per-request deadline; every attempt gets a smaller timeout than the previous one
- only retry idempotent operations, or enforce idempotency keys server-side
When to use it (and when not to)
Use a fallback chain when the business cost of “no answer” is higher than the cost of “slower or lower-quality answer.” Don’t use it just because multi-region sounds mature.
| Scenario | Recommendation |
|---|---|
| User-facing chat, search, or summarization where degraded output is acceptable | Yes: model fallback plus region fallback with a strict latency budget |
| Batch jobs that can wait minutes and be resumed | Usually no chain needed; use queue retries and dead-lettering instead |
| Requests with side effects, such as tool calls that send email or mutate records | Only if you have idempotency keys and replay protection |
| Compliance requires data residency in one region only | Region fallback may be forbidden; use model fallback within the allowed region |
| You only have one provider, one model, one region, and low availability requirements | You probably don’t need this yet |
| Model outputs must be consistent for audits or deterministic workflows | Avoid cross-model fallback unless you version and validate output schemas |
You probably don’t need this if:
- your SLA tolerates waiting for a queue retry
- your requests are internal and non-interactive
- you cannot tolerate semantic drift between models
- your biggest issue is bad prompts or quota misconfiguration, not availability
Trade-offs
Every benefit costs something operationally.
- Higher availability costs more latency. Each fallback attempt burns deadline budget and can turn a 700 ms p95 into a 2.2 s p95 during incidents.
- Better resilience costs more money. Cross-region egress, duplicate warm capacity, and health-check traffic are not free.
- Vendor flexibility costs more testing. Alternate models often differ in token limits, tool-calling behavior, JSON formatting, and safety filters.
- Faster incident recovery costs more complexity. Circuit breakers, idempotency stores, and per-attempt observability add moving parts.
- Reduced blast radius costs more lock-in work. If you abstract too aggressively around provider-specific features, you either lose useful capabilities or end up re-implementing them poorly.
The honest failure mode: fallback can hide outages long enough that you discover them through quality regressions instead of hard errors. That is better for users, but worse for diagnosis unless you emit explicit degraded-service signals.
In practice
Example 1: nginx as a region failover front door
This handles region-level failover for HTTP inference endpoints. It does not understand model semantics; it just retries another upstream on transport and selected 5xx failures.
⚠️ If your upstream operation is not idempotent,
proxy_next_upstreamcan duplicate work. Do not use this pattern for requests that trigger external side effects unless your upstream honors an idempotency key.
upstream infer_primary_secondary {
zone infer_zone 64k;
server us-east-1-infer.internal:443 max_fails=3 fail_timeout=30s;
server eu-west-1-infer.internal:443 backup;
keepalive 64;
}
server {
listen 443 ssl http2;
server_name infer.example.com;
location /v1/infer {
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_set_header Host $host;
proxy_set_header X-Request-Id $request_id;
proxy_set_header Idempotency-Key $http_idempotency_key;
proxy_connect_timeout 1s;
proxy_read_timeout 2s;
proxy_next_upstream error timeout invalid_header http_502 http_503 http_504;
proxy_next_upstream_tries 2;
proxy_pass https://infer_primary_secondary;
add_header X-Proxy-Upstream $upstream_addr always;
add_header X-Proxy-Status $upstream_status always;
}
}
What it does: sends traffic to the primary region and fails over to the backup on connect/read errors and selected 5xx statuses. Gotcha: nginx cannot decide that model A should fall back to model B based on a JSON error body; that logic belongs in your app or gateway.
Useful diagnosis command:
curl -sS -D - -o /dev/null https://infer.example.com/v1/infer
Typical output during failover:
HTTP/2 200
server: nginx
x-proxy-upstream: 10.20.4.18:443
x-proxy-status: 502, 200
That 502, 200 shape means the first upstream failed and the backup served the response.
Example 2: application-level fallback with explicit policy in Node.js
This is where model-aware logic belongs. It classifies failures, applies a deadline, and marks degraded responses.
import crypto from "node:crypto";
const chain = [
{ model: "reasoning-large", region: "us-east-1", timeoutMs: 1200 },
{ model: "reasoning-large", region: "eu-west-1", timeoutMs: 800 },
{ model: "reasoning-medium", region: "eu-west-1", timeoutMs: 500, degraded: true }
];
function shouldFailover(status, errBody, errCode, remainingMs) {
if (["ENOTFOUND", "ECONNRESET", "ECONNREFUSED", "ETIMEDOUT"].includes(errCode)) return true;
if ([502, 503, 504].includes(status)) return true;
if (status === 429) {
const retryAfter = Number(errBody?.retry_after ?? 0) * 1000;
return retryAfter === 0 || retryAfter > remainingMs;
}
return false;
}
export async function infer(payload) {
const requestId = crypto.randomUUID();
const idem = `infer-${requestId}`;
const deadline = Date.now() + 2500;
const attempts = [];
for (const candidate of chain) {
const remainingMs = deadline - Date.now();
if (remainingMs <= 0) break;
const timeoutMs = Math.min(candidate.timeoutMs, remainingMs);
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), timeoutMs);
try {
const res = await fetch(`https://${candidate.region}-infer.internal/v1/infer`, {
method: "POST",
signal: controller.signal,
headers: {
"content-type": "application/json",
"x-request-id": requestId,
"idempotency-key": idem
},
body: JSON.stringify({ model: candidate.model, input: payload })
});
clearTimeout(timer);
const body = await res.json().catch(() => ({}));
if (res.ok) {
return {
ok: true,
body,
meta: { requestId, attempts, servedModel: candidate.model, servedRegion: candidate.region, degraded: !!candidate.degraded }
};
}
attempts.push({ candidate, status: res.status, body });
if (!shouldFailover(res.status, body, null, remainingMs)) {
return { ok: false, error: body, meta: { requestId, attempts } };
}
} catch (e) {
clearTimeout(timer);
const code = e.name === "AbortError" ? "ETIMEDOUT" : e.code;
attempts.push({ candidate, error: code || String(e) });
if (!shouldFailover(null, null, code, remainingMs)) {
return { ok: false, error: { message: String(e) }, meta: { requestId, attempts } };
}
}
}
return { ok: false, error: { message: "fallback chain exhausted" }, meta: { requestId, attempts } };
}
What it does: tries three candidates in order with shrinking per-attempt timeouts and returns metadata about what actually served the request. Gotcha: if alternate models produce different JSON schemas or tool-call formats, validate the response before returning it to callers.
Example 3: quick health and DNS/TLS checks during an incident
When a region is “down,” verify whether the failure is DNS, TCP, TLS, or HTTP before editing any routing.
dig +short us-east-1-infer.internal
nc -vz us-east-1-infer.internal 443
openssl s_client -connect us-east-1-infer.internal:443 -servername us-east-1-infer.internal </dev/null
curl -sS -o /dev/null -D - --max-time 2 https://us-east-1-infer.internal/healthz
Typical failure shapes:
$ dig +short us-east-1-infer.internal
$ nc -vz us-east-1-infer.internal 443
nc: getaddrinfo for host "us-east-1-infer.internal" port 443: Name or service not known
$ nc -vz us-east-1-infer.internal 443
Connection to us-east-1-infer.internal (10.20.1.14) 443 port [tcp/https] succeeded!
$ curl -sS -o /dev/null -D - --max-time 2 https://us-east-1-infer.internal/healthz
curl: (28) Operation timed out after 2001 milliseconds with 0 bytes received
What it does: separates name resolution failures from service unavailability. Gotcha: a passing TCP connect does not mean the model service is healthy; always check an HTTP endpoint or a real synthetic inference.
Further reading
- NGINX
proxy_next_upstreamdirective documentation - The "Timeouts" and "Retries" sections of the Google SRE Book
- Martin Fowler: Circuit Breaker
- RFC 9110 HTTP Semantics
- The "Caching" and "HTTP response status codes" chapters of the MDN HTTP docs
This article was written by an AI system and published pending human review. Verify anything you intend to act on.
Have a project in mind?
Get an instant AI price estimate for it, or talk directly to our team.
One email a month on what we learn building with AI