Most AI products start the same way: one provider, one API key, a model name hard-coded in the application. It works until the day that provider has an outage, changes its prices, or ships a model update that quietly makes your answers worse. Then every one of those problems is your users' problem.
On a consumer AI platform we run, we removed that single point of failure without rewriting the product. This is the pattern, and the details that only showed up in production.
Separate the question from the model
The application should never know which model answers it. We split routing into two layers:
- The application decides the role a request needs, from the kind of task and the subscriber's tier: a quick answer, a deep search, a long document, a creative piece.
- A gateway, LiteLLM in our case, decides which model plays that role right now.
The application calls an OpenAI-compatible endpoint with a role name such as "fast" and nothing else. Changing the model behind a role becomes a configuration change, reviewed and deployed like any other, with no application release.
Give every role a fallback chain
Each role has an ordered list of deployments. On this platform the chain usually runs:
- The primary open-weight model at the main inference provider.
- A second open-weight model from a different model family, at the same provider.
- A different cloud's managed AI service as the last resort.
High-effort modes add Claude on Amazon Bedrock as a separate provider. The important property is that no single provider, model family, or region sits on every path.
Let the gateway handle failure
The gateway does the boring, essential work so the application does not have to:
- Two retries per request and a 120-second timeout.
- After three failures a deployment goes on a 30-second cooldown, which works like a circuit breaker: traffic moves down the chain instead of hammering a provider that is already struggling.
- Router state lives in Redis, so every gateway replica agrees on which deployments are cooling down.
- Spend is tracked per role, so you can see which feature costs what.
Measure the speed you are selling
Several providers sell the same model in different processing windows at different prices. It is tempting to pick the cheap one. We timed them first, with the same model, the same short answer and the median of three runs: about 2.2 seconds in the fastest window, about 8 seconds in the next, about 47 seconds in the standard window and over 3 minutes in the slowest.
On a tier whose whole promise is speed, a 47-second answer is not a cheaper version of the product; it is a different product. The fast role carries 82 to 85 percent of all calls on this platform, so that decision dominates both cost and user experience. Make it with a stopwatch, not a price sheet.
Treat reasoning as a cost line
Reasoning tokens are billed as output tokens, the most expensive kind. We measured each model instead of applying one policy everywhere:
- Some models cannot switch reasoning off at all.
- One model spent over a thousand output tokens thinking about a one-word question unless reasoning was disabled.
- Another printed its reasoning draft straight into the user-visible answer at a low reasoning setting, so that setting was banned for it.
- Where the answer rests on retrieved context with citations, switching reasoning off made no visible difference in quality and roughly halved output tokens.
Reconcile prices with the provider, not with yourself
This is the lesson we would most like other teams to skip learning the hard way. The rates per model lived in four places: the gateway config, a role catalogue, defaults in code, and the billing module. A test kept all four in agreement, and it always passed.
Then the provider changed its prices. For one model our rates were exactly double what we were actually charged; for others they were off in both directions. Every check stayed green, because all four copies were wrong in exactly the same way. The check that matched our own forecasts against real charges did not catch it either, because the most-used role was not part of that comparison.
Consistency inside your repository is not reconciliation with your supplier. We now read the provider's own price endpoint on a schedule and fail the build when it disagrees with the repository.
A checklist you can use
- The application asks for a role, never for a model name.
- Every role has at least one fallback at a different provider.
- Timeouts, retries and cooldowns live in the gateway, with shared state.
- Latency is measured per role and per processing window before choosing one.
- Reasoning settings are tested per model, including what the user actually sees.
- Prices are reconciled against the provider on a schedule, not just against your own files.
None of this needs a large platform team. It needs a gateway, a few rules written down, and the habit of measuring before deciding.