Home / Case Studies / AI platform
AI platform

An AI platform that never depends on one model provider.

A consumer generative-AI product where an in-house router picks the job, a LiteLLM gateway picks the model with automatic fallbacks across three providers, and hybrid RAG grounds answers in cited sources.

Consumer AI · Client name withheld under NDA

3

Model providers: with automatic fallback per role.

82–85%

Calls on the fast role: answered at about 2.2 s per window.

6.8 rps

Measured retrieval capacity: about 24,000 subscribers.

LiteLLMAmazon BedrockClaudePostgreSQL + pgvectorRedisKubernetesArgo CDKargoKarpenterOracle Cloud

The challenge

The product answers subscribers in chat, with paid tiers that promise different depth and speed. It depended on a single inference provider: an outage, a price change, or a model that misbehaved went straight to users. Answers also had to rest on the product's own knowledge base, with sources, rather than on whatever a model remembered.

The constraints

  • Speed is the product on the cheapest tier, so a slower but cheaper processing window was not an option.
  • Provider prices change; the cost model must follow what is actually charged.
  • Some topics need evidence: medical statements may cite research sources only, never opinion content.
  • Switching a model must never require an application release.

Decisions & tradeoffs

  • Two layers of routing. The application decides the role a request needs (task type × the subscriber's tier); LiteLLM decides which model serves that role. The application calls a role such as "fast" and never knows the model behind it, so changing a model is a configuration change.
  • Fallbacks per role across providers: the primary open-weight model, then a second open-weight model, then Oracle Cloud's Generative AI service as the last resort. Amazon Bedrock adds Claude for high-effort modes.
  • Resilience in the gateway: two retries, a 120-second timeout, and after three failures a model is put on a 30-second cooldown, like a circuit breaker, with router state in Redis.
  • Decisions by measurement: response windows were timed (2.21 s on the fast window against about 47 s on the standard one), and reasoning was switched on or off per model after live tests, because reasoning tokens are billed as output.

The implementation

Retrieval is hybrid: vector search over PostgreSQL with pgvector and bge-m3 embeddings, plus keyword (BM25) search, merged with reciprocal rank fusion and then re-ranked. Every chunk carries its source, type, date, and credibility, and answers cite them.

The gateway tracks spend per role against a price list read from the provider's own API, and tests fail the build if the price list, the role catalogue, and the cost code drift apart. The platform runs on Kubernetes with Argo CD and Kargo promotions from staging to production, and Karpenter for capacity.

Subscribers reach an API that routes each request to a role; hybrid RAG on pgvector with BM25, rank fusion and re-ranking supplies cited context; a LiteLLM gateway with Redis state sends the role to the primary open-weight models, to Amazon Bedrock for high-effort modes, and to OCI Generative AI on failure.
Generative-AI platform. Highlighted: the request path.

Outcomes

A provider outage or a misbehaving model now degrades to the next model in the chain instead of taking the product down, and model changes ship as configuration. The cheapest role carries 82 to 85 percent of calls at the speed the tier promises, and capacity testing showed retrieval sustaining 6.8 requests per second, about 24,000 subscribers before it becomes the bottleneck.

Handover & ongoing ownership

Routing rules, the price-list check, and the RAG invariants are documented in the repository and enforced by tests.

Planning something similar? Talk to an AWS partner in Armenia that has built it before.

What’s next for
your business?

Let’s talk