System Design September 3, 2026 ~9 min read

Designing a Self-Hosted AI Assistant (Part 2): The LLM Gateway

TL;DR: Every model call in Timothy’s own loop goes through one Go service, the gateway. It holds the provider credentials for those calls, picks a provider and model per request, streams the answer back in one normalized format, and writes a cost row for every attempt. Three rules shape it.

First, providers and routes are data. Adding a provider or swapping a model is a row edit, live within seconds, never a code change or a restart.

Second, failover stops at the first byte of content. Before anything reaches the client, the gateway retries and fails over freely. After that, it never switches providers: an error is relayed honestly instead of being papered over.

Third, cost is recorded, never guessed. Each attempt is priced at request time and never rewritten. An unknown price means NULL, not an estimate.

The problem

Timothy talks to many models: Anthropic and OpenAI directly, OpenAI-compatible endpoints such as Z.AI, xAI and OpenRouter, AWS Bedrock, and a local Ollama on my own hardware. Each speaks a different protocol, reports tokens differently, and fails in its own way. Chat turns, summaries, memory extraction, embeddings and mission workers all need them.

RequirementWhat it means
One credential holderThe agent loop never touches a provider key
Change without deploysAdd a provider, swap a model, reorder a fallback: no restart
One stream shapebrain’s agent loop handles one vocabulary, not four
Honest failureNo silently restarted answers
Honest costEvery attempt, success or failure, leaves a row

Challenge 1: where do providers live?

Why not configure providers in code or a config file?

OptionHow it worksWhy not (or why)
A. A Go type per providerClients and model names in codeEvery new model is a release.
B. Config file plus env varsYAML for providers, keys in the environmentA restart for every model swap, and plaintext keys in every process that inherits the environment.
C. Rows plus secret references (chosen)A providers row names a driver, base URL, default model and a credential_refEdited from the UI, hot-reloaded, and the row never holds a key.

I chose rows. A provider is a row with a driver (one of four: anthropic, openaicompat, openai-responses, bedrock), a base_url, headers, an options blob and a credential_ref. The ref is a name, never a value. One openaicompat driver serves every conforming endpoint. Drivers are thin wire adapters: they translate, never loop or route.

  • A secret store resolves refs. The secrets table maps a ref name to a backend: built-in storage (AES-256-GCM under a master key, with the ref name bound as additional data so a ciphertext copied onto another row fails to open), HashiCorp Vault, or AWS Secrets Manager. The master key, TIMOTHY_MASTER_KEY, is the only credential that lives in the environment.
  • Unresolved means unhealthy, not crashed. A ref that resolves empty marks that provider unhealthy, and routing skips it. A row whose driver fails to build is excluded alone; the rest keep serving.
  • Snapshots swap atomically. Both tables are read in one read-only REPEATABLE READ transaction into an immutable snapshot, swapped in through an atomic pointer. Reloads run on a 30-second poll and after every admin write. A failed reload keeps the last good snapshot.

Challenge 2: how does a request pick a model?

Why not let each caller name the model?

OptionHow it worksWhy not (or why)
A. Callers name a modelbrain sends a model idModel ids spread through the code; changing one is a hunt.
B. Task categoriesCallers say “coding” or “summarize”A fixed taxonomy every caller and every chain must agree on.
C. Route names in codeCode compares route == "embedding"Renaming a route silently breaks whatever looked it up.
D. Routes bound to roles (chosen)Four roles (default, summarize, embedding, vision) say which route serves what Timothy needsNames are free text; code asks “which route holds role X?”

I chose roles. A route is a named, ordered chain of {provider_id, model} entries, plus a capability (chat, embeddings or vision) and an optional role. A partial unique index allows one route per role. brain asks the gateway which route holds a role, so routes can be renamed or rebuilt without touching code, and agents can point at their own.

Route resolution from role to attempts
Figure 1: A request becomes an ordered list of attempts: hint, then sticky, then the chain in strategy order, each candidate gated on health and capability.

  • Capability is checked twice. The driver must declare it, and the model’s entry in the synced LiteLLM catalog must agree, so an embeddings-only model in a chat chain is skipped before any wire call. Images in the messages add vision to every candidate.
  • Sticky beats score. The provider that last served this session successfully on this route goes first, because staying put keeps its prompt cache warm. Sticky is a preference, never an expansion: it must already be in the chain, and editing the route resets it.
  • Strategies are per route. ordered keeps my written order. auto, price and latency score entries from declared prices and the last hour of ledger stats. Uptime multiplies the score, so a cheap provider that keeps failing sinks under every strategy.
  • Rejections are explained. If nothing passes, the no_route error lists every skipped candidate and why.

Challenge 3: when is it safe to switch providers?

If a provider fails, why not just replay the request on the next one?

The rule rests on one contract: every driver translates its wire format into the same channel of events, ending with exactly one terminal event, done or error. Drivers also normalize usage, so input tokens always exclude cache reads.

Four drivers, one event stream
Figure 2: Four wire formats in, one ordered event vocabulary out, with exactly one terminal event per stream.

OptionHow it worksWhy not (or why)
A. No failoverOne provider per requestOne outage stops every turn.
B. Replay anywhereOn any error, resend to the next providerThe client already rendered half an answer, maybe a tool call; a second model writes a different one on top.
C. SpliceContinue the partial answer on another providerTwo models writing one answer, and tool-call state cannot be handed over.
D. A bright line (chosen)Fail over only before the first content eventThe client sees one provider’s answer, or an honest error.

I chose the bright line. An attempt’s errors are held back until it produces content: a chunk, reasoning_chunk or tool event. Until then, a failed attempt is drained quietly, recorded, and the loop moves on with a failover event. After that, the stream belongs to that provider, and any error goes to the client as-is.

The failover bright line
Figure 3: Left of the line the client has seen nothing, so retry and failover are safe; right of it, the gateway only reports.

  • Not every error advances. Timeouts, 5xx, 429 and 401/403 (a bad key on this provider) move on. 400, 404, 413 and 422 stop the chain, because the next provider would reject the same request.
  • Retries are sized by position. The HTTP drivers retry 429, 5xx and network errors with jittered exponential backoff, but only once while more entries remain; the last entry gets three.
  • Exhaustion names every attempt by error code only, never raw provider text.

Challenge 4: how do you account for cost honestly?

Why not store tokens and price them later?

OptionHow it worksWhy not (or why)
A. Price at query timeStore tokens; price them when the dashboard asksA price edit rewrites history overnight.
B. Provider billing exportsReconcile from each vendor’s usage exportLate, per vendor, blind to sessions and agents.
C. Price at request time (chosen)Cost computed when usage arrives and frozen in the rowThe row says what was known then; unknown stays NULL.

I chose to freeze cost at request time. Every attempt, failed ones included, leaves one cost_ledger row: provider, model, route, agent, purpose, session, tokens, latency, status and cost. The done event carries the ledger_id and cost back to brain. A failed insert is logged, never raised: accounting must not break serving.

Lifecycle of a ledger row
Figure 4: Prices are resolved before the call, cost is fixed when usage arrives, and the same frozen rows feed stickiness, scoring and budgets.

  • Prices come from two places. An operator’s per-model override wins; otherwise the synced catalog. If neither has the model, cost is NULL, with a warning and an unpriced-usage counter, so the gap is visible instead of silently zero.
  • Cache writes are split by tier. One-hour cache writes are stored apart from five-minute ones because they bill higher. With no declared one-hour price, the cost is NULL rather than a guessed multiple.
  • Currency is stored, never converted. FX rates exist for display only.
  • The same rows steer routing. Stickiness, scored strategies, the usage dashboard and daily and monthly budgets all read the ledger.

What broke

Some reasoning-class OpenAI models returned a 200 and an empty stream over chat/completions on turns where a tool call was mandatory. Counting that as success is doubly wrong: the turn looks fine, and stickiness, which reads status='ok' rows, pins the session to the provider that returned nothing. Now a contentless stream is booked empty_output and advances the chain, and a dedicated openai-responses driver serves those models. Success means content, not a status code.

What’s next

Part 3 follows the stream into brain, where every session is an append-only event log with two projections.

// end of article — process exited with code 0

// share:Twitter / XLinkedIn
// llm:View as Markdown

// comments

loading...