Skip to main content

Handle LLM API outages: retries, fallback and graceful degradation

Keep SaaS AI features reliable during provider errors with error classification, bounded retries, circuit breakers, clear degraded behavior and tested provider failover.

In this guide

How should a SaaS application handle an LLM provider outage?

Separate temporary provider trouble from errors that will not recover by retrying, then apply a bounded recovery plan. Respect provider retry guidance, use a small retry budget with backoff, open a circuit when failures persist, and return a clear degraded response when the AI feature cannot safely run. A provider outage should not trigger an unbounded agent loop, duplicate customer actions or a silent switch that changes data handling or answer quality.

Classify errors before deciding whether to retry

Distinguish temporary overload, network timeout and rate limits from invalid requests, missing permissions, exhausted credits and spend limits. Retryable conditions may merit a bounded delay; authentication, schema and billing errors need correction, not repeated calls. Read documented status, error code and Retry-After information for your provider and SDK version instead of treating every non-success as a transient outage.

Bound attempts and respect retry-after guidance

Set a maximum attempts count, total time deadline and per-tenant retry budget. When the provider supplies a Retry-After delay, wait at least that long; otherwise use exponential backoff with jitter and avoid synchronized retry storms. Coordinate retries at one layer in the call stack, and account for automatic SDK retries so your application does not multiply them unknowingly.

Make uncertain outcomes safe to resume

A timeout may occur after a provider or downstream tool completed work. Track request state and use idempotency keys for side effects so a retry does not send duplicate messages, charge an account or modify the same record twice. Do not retry invalid, unauthorized or explicitly non-retryable work merely because the UI displayed an error.

LLM provider outage response matrix
Error classRetry? and limitUser-visible behaviorFallback/stop triggerOwner and test
Temporary overload
Rate limit
Auth/configuration error

How do circuit breakers and degraded modes protect customers?

Open a circuit when provider failures persist

Measure timeouts, overload responses and dependency health. Stop sending new traffic for a defined window after a failure threshold, then probe recovery with limited requests. Keep the circuit scoped to the failing provider, model route or feature so one outage does not unnecessarily disable unrelated SaaS functions.

Define a useful and honest degraded experience

When safe, allow a user to save a draft, browse already-authorized source material, retry later or complete the task manually. Clearly label stale or partial results and do not present a fallback answer as current, verified or equivalent when it is not. If a safe partial result is unavailable, return a clear error and a next step rather than inventing an answer.

Control queues and shed load before a retry cascade

Limit queue depth, concurrency and waiting time; expire abandoned tasks; and prioritize work according to a documented customer need. Reject or defer excess work cheaply when the provider is overloaded instead of keeping workers occupied with requests that cannot succeed. Monitor retry ratio and queue age as early signs of a growing incident.

When is a second AI provider a safe fallback?

Review data terms and location before moving requests

Confirm the alternate provider, endpoint, region, retention and subprocessors are approved for the data and customer promise. Do not send a customer's prompt to a second service during an outage unless the service's processing and transfer terms allow it. If approval is absent, degrade or pause the AI feature instead.

Exercise failover and recovery without production side effects

Run scheduled tests using synthetic or approved traffic. Verify the circuit opens, the alternate route works, usage remains within budget, customers see accurate status and recovery does not switch traffic back and forth rapidly. Document how to disable the fallback and how to reconcile work that was in flight during the transition.

LLM API outage FAQs