Everybody building on an LLM eventually writes a fallback. The provider goes down, you switch to another one, the feature stays up. It is the obvious thing to build and it is worth building.
But an outage is only one way a call fails. Another is a provider that is healthy, answers in normal time, and returns nothing usable because it decided not to answer: a content filter, a safety block. That is a refusal, and it needs different handling from an outage.
We work through this on Snark, an open-source API that generates AI-written humor responses across Groq, Gemini and Claude. Everything below describes what that code does.
01A refusal is not an outage
If one provider is down, moving to the next one is the right answer: the outage belongs to the provider. A refusal is triggered by the request, so another provider may or may not accept it, and you pay a full extra round trip to find out.
A blanket except Exception: try_next_provider() also hides the difference. A refusal and a dead socket look identical in your logs, so you cannot tell which one you are tuning for. The first fix is small: give the two failures different types.
class ProviderError(Exception):
"""Raised when an AI provider fails to generate a response."""
class ContentFilterError(ProviderError):
"""Raised when a provider blocks output due to content filtering."""ContentFilterError is a subclass, so a caller that does not care catches both with one handler, while a caller that does care can tell them apart. Each provider adapter translates its own signal into it: Groq's finish_reason of content_filter, a Gemini blocked prompt, safety finish reason or empty response, and a Claude bad-request error whose message contains "content filtering" or "blocked".
02What happens on a refusal
For a normal (non-streamed) request, a content-filter error on the primary provider does not send the request straight to the next provider. The service retries the same provider first, with a softer system prompt and a lower temperature:
softened_system = (
system_prompt
+ "\n\nIMPORTANT: Keep it light, playful, and safe for all audiences. "
"Avoid anything offensive, mean-spirited, or inappropriate."
)
# …
temperature=max(temperature - 0.2, 0.3)The temperature drops by 0.2 with a floor of 0.3. Only if this softened retry also fails does the service walk the fallback list, and it sends the original, unsoftened prompt to those providers.
Primary provider
The default provider from settings (groq unless configured otherwise) gets the request first.
Only on a content filter
Softened retry
Same provider, softened system prompt, temperature lowered by 0.2 with a floor of 0.3.
If that fails, or on a provider error
Fallbacks in order
Each available provider in the configured order, sent the original prompt. First success wins.
If every provider fails
Error raised
A ProviderError is raised and a failed event is recorded with an error code.
03The fallback list
Order lives in settings, read from the environment, rather than being hard-coded in the request path:
AI_DEFAULT_PROVIDER=groq
AI_PROVIDER_FALLBACK_ORDER=groq,gemini,claudeTwo details in the registry matter more than they look:
- Availability is checked before a provider joins the list.
get_fallbacks()only returns providers whoseis_available()is true; for Groq that means the package is installed andGROQ_API_KEYis set. A provider with no API key is not a fallback, it is a guaranteed second failure. - The primary is excluded from its own list. The configured order starts with the default provider, so without the
excludeargument the primary would be retried as its own fallback.
Provider instances are created lazily and cached for the life of the process.
04Streaming behaves differently
Streamed responses take a separate path, and it is deliberately simpler. There is no softened retry: on a failure the service moves to the next provider in the list. And once any text has already been sent to the client, it cannot switch providers cleanly, so it logs the mid-stream failure and stops instead of stitching two providers' output together.
Streams also bypass the response cache, because a stream cannot be replayed from cache the way a finished response can.
05Record whether failover happened
Every generation writes an event with the fields success, fell_back, content_filtered, streamed, error_code and error_detail, and a Prometheus counter, snark_generations_total, is labelled by provider, success, fell_back and content_filtered.
fell_back is true when a non-primary provider served the request. Without it, a failover layer that works looks identical to one that is never used: users get answers, error rate stays flat, and you cannot see how much traffic is landing on your second or third choice. content_filtered stays true even when a fallback provider ends up serving the request, so a refusal that was recovered from is still counted. That separates refusals from outages, so the two can be measured on their own.
06What it does not do
- No circuit breaker. If the primary is down, every request tries it first and waits before falling back.
- Static order. The fallback order does not adapt to cost or current latency; it is whatever the setting says.
The failover path is in snark/wit/services.py, and the provider abstraction and registry are in snark/wit/providers/. Snark is open source under AGPL-3.0, so if we have described something wrong, the code is there to check.
Building something that needs to stay up?
We build and harden AI systems, blockchain infrastructure and the web layer around them.
Start a project