AI Engineering
Multi-tenant agentic support platform
A multi-tenant conversational AI platform that routes each request through a contextual bandit, runs a parallel perception pipeline, retrieves per-tenant knowledge with hybrid search, and enforces policy and budgets on every action.
Client work. The code is private, so this page covers the engineering approach only.
Problem
A conversational AI product that answers on behalf of many separate companies has to solve three problems at once, and they pull in different directions. It has to keep every tenant's data provably isolated, it has to spend model budget carefully because the cheap path and the expensive path can differ by an order of magnitude per request, and it has to stay safe and on-policy on inputs it has never seen. Treating the LLM as a single black box endpoint fails all three: there is no place to enforce tenant boundaries, no signal to decide how much reasoning a message needs, and no gate between the model deciding to act and the action happening.
The other constraint is operational. Traffic is bursty and multilingual, a fraction of it is abusive or in crisis, and a fraction needs a human. The system needs to read intent, sentiment, language, and frustration fast enough to route in real time, escalate the cases that warrant it, and fall back to a safe default when a component is down, without dropping the conversation.
The platform is built as a fleet of focused services behind a gateway so each of these concerns lives in one place with a clear boundary, and changing routing does not touch tenant isolation.
How it works
A FastAPI gateway fronts a set of independent services (perception, routing, runtime, knowledge, memory, governance, provisioning, compliance) that share a small set of internal libraries for models, schemas, and telemetry. Requests carry a tenant identity from the first hop, and that identity is what scopes data, budget, and policy for the rest of the request.
Each request is classified before it reaches a model, then routed to a specific model and reasoning depth, then checked against tenant policy before any tool runs. Retrieval, memory, and observability hang off that spine as separate services.
- Dual-path authentication at the gateway: an internal HS256 token for platform users and Keycloak-issued RS256 tokens for SSO, both fully signature- and expiry-verified before the tenant is resolved.
- Tenant isolation is enforced at the database layer. Middleware resolves the tenant, verifies its lifecycle state, and sets a per-request PostgreSQL session variable so Row-Level Security scopes every query even if the application forgets a WHERE clause.
- Model and reasoning-strategy selection runs as a contextual multi-armed bandit using Beta-Bernoulli Thompson Sampling, keyed on a context bucket of agent, task complexity, and intent, with a warm-up phase that falls back to a global prior until a context has enough observations.
- A complexity classifier scores each prompt across lexical, task-type, ambiguity, context-dependency, and domain signals and maps the weighted score to one of three reasoning strategies (direct call, lightweight structured reasoning, full chain-of-thought) so cheap questions do not pay for deep reasoning.
- Every action is checked against an Open Policy Agent ruleset written in Rego that is fail-closed by default: it decides tool allowlisting, role and action-class permissions, forced human-in-the-loop patterns, and per-tier token budget enforcement, and returns a machine-readable denial reason.
- Per-tenant knowledge retrieval is hybrid: dense vectors in Qdrant and sparse BM25S results are merged with Reciprocal Rank Fusion, normalized, and capped at the top few, with a knowledge-gap flag raised when every score falls below threshold.
- Long-running and multi-step processes (human approval, escalation, data erasure, tenant offboarding, memory consolidation) run as Temporal workflows, and nightly per-tenant background cycles run as Celery tasks under a distributed Redis lock with a hard time budget.
Hard parts
- Real-time perception without serial latency: the six analysis stages (safety, intent, entities, sentiment, emotion, language) fan out in parallel with asyncio.gather, the frustration score is computed from the sentiment and emotion outputs, and any stage that raises is replaced with a safe default and recorded as a partial failure so one slow model never stalls or blocks the whole read. The pipeline targets sub-500ms and truncates long inputs per stage, giving safety a larger window than the rest.
- Cold-start routing: a fresh context has no reward history, so the bandit uses an uninformative prior, a higher exploration rate during warm-up, and a low-confidence flag until it crosses an observation threshold, then decays exploration over time toward a floor so the router keeps learning without thrashing settled contexts.
- Provable tenant boundaries: enforcing scoping in application code is easy to bypass by accident, so isolation is pushed into the database with Row-Level Security driven by a request-scoped session variable, and pre-principal paths like login are handled explicitly so authentication can read what it needs without leaking cross-tenant data.
- Safe autonomy: the model does not decide on its own when to act; policy sits between intent and execution. The Rego layer can force copilot mode for sensitive action patterns, block tools outside a tenant's allowlist, and stop work when a budget is exhausted, all with a default-deny posture.
- Bounded background work at scale: the nightly per-tenant cycle orders tenants by plan tier, holds a distributed Redis lock per tenant to prevent overlap, enforces a total wall-clock budget that skips the remainder when spent, and runs phases sequentially with early termination and per-phase metrics so a failure produces a partial report that says what ran.
What it does
The design turns three competing requirements into three independent, testable mechanisms. Isolation, cost control, and safety each have a single enforcement point that can be verified on its own, and the routing layer improves its model and reasoning choices from observed reward.
- Per-request reasoning depth is chosen from prompt signals, so simple lookups avoid the cost of full chain-of-thought while hard questions still get it.
- Model-strategy selection adapts per context through Thompson Sampling, with exploration that decays as evidence accumulates.
- Tenant data isolation is enforced at the database layer and holds even when application code omits an explicit filter.
- The perception read stays within its latency target and returns a usable result even when individual stages fail.
- Policy decisions are fail-closed and auditable, with an explicit denial reason on every blocked action.
Artifacts
The source lives in a private repository and is not public. This entry describes the engineering approach in general terms without exposing product identifiers, internal endpoints, or proprietary configuration.