# Designing a Self-Hosted AI Assistant (Part 2): The LLM Gateway

> Source: https://www.sumonselim.com/timothy-part2-the-llm-gateway/
> Author: Muhammad Sumon Molla Selim
> Published: 2026-09-03
> Tag: System Design
> Summary: Every model call in Timothy's own agent loop goes through one gateway that holds the provider keys. Providers are Postgres rows, routes are bound to roles and filtered by capability, every driver speaks one event vocabulary, and failover stops the moment content reaches the client. Every attempt leaves a ledger row priced at request time, or NULL when the price is unknown.
>
> Series: Designing a Self-Hosted AI Assistant, part 2 of 7
> 1. [The Architecture](https://www.sumonselim.com/timothy-part1-the-architecture.md)
> 2. The LLM Gateway (this part)
> 3. [Sessions as an Event Log](https://www.sumonselim.com/timothy-part3-sessions-as-an-event-log.md)
> 4. [Never Lose a Turn](https://www.sumonselim.com/timothy-part4-never-lose-a-turn.md)
> 5. [Memory and Knowledge](https://www.sumonselim.com/timothy-part5-memory-and-knowledge.md)
> 6. [Letting the Model Act Safely](https://www.sumonselim.com/timothy-part6-letting-the-model-act-safely.md)
> 7. [Missions](https://www.sumonselim.com/timothy-part7-missions.md)

**TL;DR:** Every model call in Timothy's own loop goes through one Go service, the **gateway**. It holds the provider credentials for those calls, picks a provider and model per request, streams the answer back in one normalized format, and writes a cost row for every attempt. Three rules shape it.

First, **providers and routes are data**. Adding a provider or swapping a model is a row edit, live within seconds, never a code change or a restart.

Second, **failover stops at the first byte of content**. Before anything reaches the client, the gateway retries and fails over freely. After that, it never switches providers: an error is relayed honestly instead of being papered over.

Third, **cost is recorded, never guessed**. Each attempt is priced at request time and never rewritten. An unknown price means `NULL`, not an estimate.

## The problem

Timothy talks to many models: Anthropic and OpenAI directly, OpenAI-compatible endpoints such as Z.AI, xAI and OpenRouter, AWS Bedrock, and a local Ollama on my own hardware. Each speaks a different protocol, reports tokens differently, and fails in its own way. Chat turns, summaries, memory extraction, embeddings and mission workers all need them.

| Requirement | What it means |
| :-- | :-- |
| One credential holder | The agent loop never touches a provider key |
| Change without deploys | Add a provider, swap a model, reorder a fallback: no restart |
| One stream shape | brain's agent loop handles one vocabulary, not four |
| Honest failure | No silently restarted answers |
| Honest cost | Every attempt, success or failure, leaves a row |

## Challenge 1: where do providers live?

**Why not configure providers in code or a config file?**

| Option | How it works | Why not (or why) |
| :-- | :-- | :-- |
| **A. A Go type per provider** | Clients and model names in code | Every new model is a release. |
| **B. Config file plus env vars** | YAML for providers, keys in the environment | A restart for every model swap, and plaintext keys in every process that inherits the environment. |
| **C. Rows plus secret references (chosen)** | A `providers` row names a driver, base URL, default model and a `credential_ref` | Edited from the UI, hot-reloaded, and the row never holds a key. |

**I chose rows.** A provider is a row with a `driver` (one of four: `anthropic`, `openaicompat`, `openai-responses`, `bedrock`), a `base_url`, headers, an `options` blob and a `credential_ref`. The ref is a name, never a value. One `openaicompat` driver serves every conforming endpoint. Drivers are thin wire adapters: they translate, never loop or route.

- **A secret store resolves refs.** The `secrets` table maps a ref name to a backend: built-in storage (AES-256-GCM under a master key, with the ref name bound as additional data so a ciphertext copied onto another row fails to open), HashiCorp Vault, or AWS Secrets Manager. The master key, `TIMOTHY_MASTER_KEY`, is the only credential that lives in the environment.
- **Unresolved means unhealthy, not crashed.** A ref that resolves empty marks that provider unhealthy, and routing skips it. A row whose driver fails to build is excluded alone; the rest keep serving.
- **Snapshots swap atomically.** Both tables are read in one read-only `REPEATABLE READ` transaction into an immutable snapshot, swapped in through an atomic pointer. Reloads run on a 30-second poll and after every admin write. A failed reload keeps the last good snapshot.

## Challenge 2: how does a request pick a model?

**Why not let each caller name the model?**

| Option | How it works | Why not (or why) |
| :-- | :-- | :-- |
| **A. Callers name a model** | brain sends a model id | Model ids spread through the code; changing one is a hunt. |
| **B. Task categories** | Callers say "coding" or "summarize" | A fixed taxonomy every caller and every chain must agree on. |
| **C. Route names in code** | Code compares `route == "embedding"` | Renaming a route silently breaks whatever looked it up. |
| **D. Routes bound to roles (chosen)** | Four roles (`default`, `summarize`, `embedding`, `vision`) say which route serves what Timothy needs | Names are free text; code asks "which route holds role X?" |

**I chose roles.** A route is a named, ordered chain of `{provider_id, model}` entries, plus a `capability` (`chat`, `embeddings` or `vision`) and an optional `role`. A partial unique index allows one route per role. brain asks the gateway which route holds a role, so routes can be renamed or rebuilt without touching code, and agents can point at their own.

![Route resolution from role to attempts](https://www.sumonselim.com/images/articles/timothy/p2-route-resolution.svg "Figure 1: A request becomes an ordered list of attempts: hint, then sticky, then the chain in strategy order, each candidate gated on health and capability.")

- **Capability is checked twice.** The driver must declare it, and the model's entry in the synced LiteLLM catalog must agree, so an embeddings-only model in a chat chain is skipped before any wire call. Images in the messages add `vision` to every candidate.
- **Sticky beats score.** The provider that last served this session successfully on this route goes first, because staying put keeps its prompt cache warm. Sticky is a preference, never an expansion: it must already be in the chain, and editing the route resets it.
- **Strategies are per route.** `ordered` keeps my written order. `auto`, `price` and `latency` score entries from declared prices and the last hour of ledger stats. Uptime multiplies the score, so a cheap provider that keeps failing sinks under every strategy.
- **Rejections are explained.** If nothing passes, the `no_route` error lists every skipped candidate and why.

## Challenge 3: when is it safe to switch providers?

**If a provider fails, why not just replay the request on the next one?**

The rule rests on one contract: every driver translates its wire format into the same channel of events, ending with exactly one terminal event, `done` or `error`. Drivers also normalize usage, so input tokens always exclude cache reads.

![Four drivers, one event stream](https://www.sumonselim.com/images/articles/timothy/p2-event-stream.svg "Figure 2: Four wire formats in, one ordered event vocabulary out, with exactly one terminal event per stream.")

| Option | How it works | Why not (or why) |
| :-- | :-- | :-- |
| **A. No failover** | One provider per request | One outage stops every turn. |
| **B. Replay anywhere** | On any error, resend to the next provider | The client already rendered half an answer, maybe a tool call; a second model writes a different one on top. |
| **C. Splice** | Continue the partial answer on another provider | Two models writing one answer, and tool-call state cannot be handed over. |
| **D. A bright line (chosen)** | Fail over only before the first content event | The client sees one provider's answer, or an honest error. |

**I chose the bright line.** An attempt's errors are held back until it produces content: a `chunk`, `reasoning_chunk` or tool event. Until then, a failed attempt is drained quietly, recorded, and the loop moves on with a `failover` event. After that, the stream belongs to that provider, and any error goes to the client as-is.

![The failover bright line](https://www.sumonselim.com/images/articles/timothy/p2-failover-bright-line.svg "Figure 3: Left of the line the client has seen nothing, so retry and failover are safe; right of it, the gateway only reports.")

- **Not every error advances.** Timeouts, 5xx, 429 and 401/403 (a bad key on this provider) move on. 400, 404, 413 and 422 stop the chain, because the next provider would reject the same request.
- **Retries are sized by position.** The HTTP drivers retry 429, 5xx and network errors with jittered exponential backoff, but only once while more entries remain; the last entry gets three.
- **Exhaustion names every attempt** by error code only, never raw provider text.

## Challenge 4: how do you account for cost honestly?

**Why not store tokens and price them later?**

| Option | How it works | Why not (or why) |
| :-- | :-- | :-- |
| **A. Price at query time** | Store tokens; price them when the dashboard asks | A price edit rewrites history overnight. |
| **B. Provider billing exports** | Reconcile from each vendor's usage export | Late, per vendor, blind to sessions and agents. |
| **C. Price at request time (chosen)** | Cost computed when `usage` arrives and frozen in the row | The row says what was known then; unknown stays `NULL`. |

**I chose to freeze cost at request time.** Every attempt, failed ones included, leaves one `cost_ledger` row: provider, model, route, agent, purpose, session, tokens, latency, status and cost. The `done` event carries the `ledger_id` and cost back to brain. A failed insert is logged, never raised: accounting must not break serving.

![Lifecycle of a ledger row](https://www.sumonselim.com/images/articles/timothy/p2-ledger-row.svg "Figure 4: Prices are resolved before the call, cost is fixed when usage arrives, and the same frozen rows feed stickiness, scoring and budgets.")

- **Prices come from two places.** An operator's per-model override wins; otherwise the synced catalog. If neither has the model, `cost` is `NULL`, with a warning and an unpriced-usage counter, so the gap is visible instead of silently zero.
- **Cache writes are split by tier.** One-hour cache writes are stored apart from five-minute ones because they bill higher. With no declared one-hour price, the cost is `NULL` rather than a guessed multiple.
- **Currency is stored, never converted.** FX rates exist for display only.
- **The same rows steer routing.** Stickiness, scored strategies, the usage dashboard and daily and monthly budgets all read the ledger.

## What broke

Some reasoning-class OpenAI models returned a `200` and an empty stream over chat/completions on turns where a tool call was mandatory. Counting that as success is doubly wrong: the turn looks fine, and stickiness, which reads `status='ok'` rows, pins the session to the provider that returned nothing. Now a contentless stream is booked `empty_output` and advances the chain, and a dedicated `openai-responses` driver serves those models. Success means content, not a status code.

## What's next

Part 3 follows the stream into brain, where every session is an append-only event log with two projections.
