Back to projects

AI Engineering

Multi-provider LLM service with failover

An async content service that routes each request to a suitable model, fails over between providers, keeps each workspace’s knowledge separate, and screens content on the way in and out.

Client work. The code is private, so this page covers the engineering approach only.

PythonFastAPIasyncioVector searchBackground jobsPrometheusDocker

Problem

The service takes user-written content and returns improved text, generated media and answers grounded in that user’s own material. Relying on one model or provider makes that fragile: a provider rate-limits or goes down, a paid media API runs over budget, and the whole request fails.

It also cannot trust its inputs or outputs. User content can carry personal data, prompt injection or harmful material, and generated media can come back unsafe.

How it works

An async FastAPI application. Heavy work runs as background jobs: the API validates and queues a request, a job processor runs the pipeline, and the client polls for the result.

Each text request is classified by what it needs (creative writing, tone preservation, speed, general use). That picks a provider order and settings, and the router walks down that order until a call succeeds. Media generation goes through a second router with the same shape.

  • Retries with backoff handle rate limits on one provider; errors that retrying will not fix move the request to the next provider.
  • A circuit breaker per provider stops calls to a provider that keeps failing, then lets a single test call through before trusting it again.
  • Knowledge is chunked, embedded and deduplicated, and every chunk is tagged with its workspace, so retrieval never returns another workspace’s content.
  • Media jobs are spread across providers by weight, and a spending limit switches traffic to cheaper inference once it is reached.
  • Moderation, personal-data detection and prompt-injection checks run in parallel on input, and generated images and video are screened before they are returned.

Hard parts

  • Keeping requests alive when a provider fails: retrying and falling over to another provider are separate steps, and the breaker keeps the router from repeatedly calling a provider that is already down.
  • Recovering jobs left behind by a crashed worker, without two workers picking up the same job: recovery runs under a lease that only one worker can hold at a time.
  • Choosing which way to fail: input that fails a safety check is rejected, while a moderation outage does not block content that was already generated.

What it does

The service reports per-provider request counts, latency, token use and safety decisions as Prometheus metrics, so fallbacks and load balancing can be seen as they happen. Requests carry a correlation ID through the async pipeline for tracing. It runs in Docker and ships through CI/CD.

Artifacts

Built for a client under contract. The code is private, so this page describes the engineering approach only.

Back to projects