3
Model providers: with automatic fallback per role.
A consumer generative-AI product where an in-house router picks the job, a LiteLLM gateway picks the model with automatic fallbacks across three providers, and hybrid RAG grounds answers in cited sources.
Consumer AI · Client name withheld under NDA
Model providers: with automatic fallback per role.
Calls on the fast role: answered at about 2.2 s per window.
Measured retrieval capacity: about 24,000 subscribers.
The product answers subscribers in chat, with paid tiers that promise different depth and speed. It depended on a single inference provider: an outage, a price change, or a model that misbehaved went straight to users. Answers also had to rest on the product's own knowledge base, with sources, rather than on whatever a model remembered.
Retrieval is hybrid: vector search over PostgreSQL with pgvector and bge-m3 embeddings, plus keyword (BM25) search, merged with reciprocal rank fusion and then re-ranked. Every chunk carries its source, type, date, and credibility, and answers cite them.
The gateway tracks spend per role against a price list read from the provider's own API, and tests fail the build if the price list, the role catalogue, and the cost code drift apart. The platform runs on Kubernetes with Argo CD and Kargo promotions from staging to production, and Karpenter for capacity.

A provider outage or a misbehaving model now degrades to the next model in the chain instead of taking the product down, and model changes ship as configuration. The cheapest role carries 82 to 85 percent of calls at the speed the tier promises, and capacity testing showed retrieval sustaining 6.8 requests per second, about 24,000 subscribers before it becomes the bottleneck.
Routing rules, the price-list check, and the RAG invariants are documented in the repository and enforced by tests.
Planning something similar? Talk to an AWS partner in Armenia that has built it before.