All writing
Engineering

Your LLM provider didn't go down. It refused.

Most LLM failover is built for outages. A content-filter refusal is a different failure, and switching providers is not the right first move.

Quorix Engineering

Quorix Technologies

Published
Reading time
3 min read

Everybody building on an LLM eventually writes a fallback. The provider goes down, you switch to another one, the feature stays up. It is the obvious thing to build and it is worth building.

But an outage is only one way a call fails. Another is a provider that is healthy, answers in normal time, and returns nothing usable because it decided not to answer: a content filter, a safety block. That is a refusal, and it needs different handling from an outage.

We work through this on Snark, an open-source API that generates AI-written humor responses across Groq, Gemini and Claude. Everything below describes what that code does.

snark — zsh
~/snark % curl http://localhost:8100/v1/wit/proverb/
{"response": "He who checks phone during conversation is lost in a sea of distraction, says the Scroll of Digital Wisdom.", "persona": "The Ancient Sage", "cached": false}

01A refusal is not an outage

If one provider is down, moving to the next one is the right answer: the outage belongs to the provider. A refusal is triggered by the request, so another provider may or may not accept it, and you pay a full extra round trip to find out.

A blanket except Exception: try_next_provider() also hides the difference. A refusal and a dead socket look identical in your logs, so you cannot tell which one you are tuning for. The first fix is small: give the two failures different types.

snark/wit/providers/base.py
class ProviderError(Exception):
    """Raised when an AI provider fails to generate a response."""

class ContentFilterError(ProviderError):
    """Raised when a provider blocks output due to content filtering."""

ContentFilterError is a subclass, so a caller that does not care catches both with one handler, while a caller that does care can tell them apart. Each provider adapter translates its own signal into it: Groq's finish_reason of content_filter, a Gemini blocked prompt, safety finish reason or empty response, and a Claude bad-request error whose message contains "content filtering" or "blocked".

02What happens on a refusal

For a normal (non-streamed) request, a content-filter error on the primary provider does not send the request straight to the next provider. The service retries the same provider first, with a softer system prompt and a lower temperature:

snark/wit/services.py
softened_system = (
    system_prompt
    + "\n\nIMPORTANT: Keep it light, playful, and safe for all audiences. "
    "Avoid anything offensive, mean-spirited, or inappropriate."
)
# …
temperature=max(temperature - 0.2, 0.3)

The temperature drops by 0.2 with a floor of 0.3. Only if this softened retry also fails does the service walk the fallback list, and it sends the original, unsoftened prompt to those providers.

  1. Primary provider

    The default provider from settings (groq unless configured otherwise) gets the request first.

  2. Only on a content filter

    Softened retry

    Same provider, softened system prompt, temperature lowered by 0.2 with a floor of 0.3.

  3. If that fails, or on a provider error

    Fallbacks in order

    Each available provider in the configured order, sent the original prompt. First success wins.

  4. If every provider fails

    Error raised

    A ProviderError is raised and a failed event is recorded with an error code.

03The fallback list

Order lives in settings, read from the environment, rather than being hard-coded in the request path:

.env.example
AI_DEFAULT_PROVIDER=groq
AI_PROVIDER_FALLBACK_ORDER=groq,gemini,claude

Two details in the registry matter more than they look:

  • Availability is checked before a provider joins the list. get_fallbacks() only returns providers whose is_available() is true; for Groq that means the package is installed and GROQ_API_KEY is set. A provider with no API key is not a fallback, it is a guaranteed second failure.
  • The primary is excluded from its own list. The configured order starts with the default provider, so without the exclude argument the primary would be retried as its own fallback.

Provider instances are created lazily and cached for the life of the process.

04Streaming behaves differently

Streamed responses take a separate path, and it is deliberately simpler. There is no softened retry: on a failure the service moves to the next provider in the list. And once any text has already been sent to the client, it cannot switch providers cleanly, so it logs the mid-stream failure and stops instead of stitching two providers' output together.

Streams also bypass the response cache, because a stream cannot be replayed from cache the way a finished response can.

05Record whether failover happened

Every generation writes an event with the fields success, fell_back, content_filtered, streamed, error_code and error_detail, and a Prometheus counter, snark_generations_total, is labelled by provider, success, fell_back and content_filtered.

fell_back is true when a non-primary provider served the request. Without it, a failover layer that works looks identical to one that is never used: users get answers, error rate stays flat, and you cannot see how much traffic is landing on your second or third choice. content_filtered stays true even when a fallback provider ends up serving the request, so a refusal that was recovered from is still counted. That separates refusals from outages, so the two can be measured on their own.

06What it does not do

  • No circuit breaker. If the primary is down, every request tries it first and waits before falling back.
  • Static order. The fallback order does not adapt to cost or current latency; it is whatever the setting says.

The failover path is in snark/wit/services.py, and the provider abstraction and registry are in snark/wit/providers/. Snark is open source under AGPL-3.0, so if we have described something wrong, the code is there to check.

Open sourcegithub.com/PramodTKodag/snarkRead the code this article describes.

Building something that needs to stay up?

We build and harden AI systems, blockchain infrastructure and the web layer around them.

Start a project