TL;DR: Every model call in Timothy’s own loop goes through one Go service, the gateway. It holds the provider credentials for those calls, picks a provider and model per request, streams the answer back in one normalized format, and writes a cost row for every attempt. Three rules shape it.
First, providers and routes are data. Adding a provider or swapping a model is a row edit, live within seconds, never a code change or a restart.
Second, failover stops at the first byte of content. Before anything reaches the client, the gateway retries and fails over freely. After that, it never switches providers: an error is relayed honestly instead of being papered over.
Third, cost is recorded, never guessed. Each attempt is priced at request time and never rewritten. An unknown price means NULL, not an estimate.
The problem
Timothy talks to many models: Anthropic and OpenAI directly, OpenAI-compatible endpoints such as Z.AI, xAI and OpenRouter, AWS Bedrock, and a local Ollama on my own hardware. Each speaks a different protocol, reports tokens differently, and fails in its own way. Chat turns, summaries, memory extraction, embeddings and mission workers all need them.
| Requirement | What it means |
|---|---|
| One credential holder | The agent loop never touches a provider key |
| Change without deploys | Add a provider, swap a model, reorder a fallback: no restart |
| One stream shape | brain’s agent loop handles one vocabulary, not four |
| Honest failure | No silently restarted answers |
| Honest cost | Every attempt, success or failure, leaves a row |
Challenge 1: where do providers live?
Why not configure providers in code or a config file?
| Option | How it works | Why not (or why) |
|---|---|---|
| A. A Go type per provider | Clients and model names in code | Every new model is a release. |
| B. Config file plus env vars | YAML for providers, keys in the environment | A restart for every model swap, and plaintext keys in every process that inherits the environment. |
| C. Rows plus secret references (chosen) | A providers row names a driver, base URL, default model and a credential_ref | Edited from the UI, hot-reloaded, and the row never holds a key. |
I chose rows. A provider is a row with a driver (one of four: anthropic, openaicompat, openai-responses, bedrock), a base_url, headers, an options blob and a credential_ref. The ref is a name, never a value. One openaicompat driver serves every conforming endpoint. Drivers are thin wire adapters: they translate, never loop or route.
- A secret store resolves refs. The
secretstable maps a ref name to a backend: built-in storage (AES-256-GCM under a master key, with the ref name bound as additional data so a ciphertext copied onto another row fails to open), HashiCorp Vault, or AWS Secrets Manager. The master key,TIMOTHY_MASTER_KEY, is the only credential that lives in the environment. - Unresolved means unhealthy, not crashed. A ref that resolves empty marks that provider unhealthy, and routing skips it. A row whose driver fails to build is excluded alone; the rest keep serving.
- Snapshots swap atomically. Both tables are read in one read-only
REPEATABLE READtransaction into an immutable snapshot, swapped in through an atomic pointer. Reloads run on a 30-second poll and after every admin write. A failed reload keeps the last good snapshot.
Challenge 2: how does a request pick a model?
Why not let each caller name the model?
| Option | How it works | Why not (or why) |
|---|---|---|
| A. Callers name a model | brain sends a model id | Model ids spread through the code; changing one is a hunt. |
| B. Task categories | Callers say “coding” or “summarize” | A fixed taxonomy every caller and every chain must agree on. |
| C. Route names in code | Code compares route == "embedding" | Renaming a route silently breaks whatever looked it up. |
| D. Routes bound to roles (chosen) | Four roles (default, summarize, embedding, vision) say which route serves what Timothy needs | Names are free text; code asks “which route holds role X?” |
I chose roles. A route is a named, ordered chain of {provider_id, model} entries, plus a capability (chat, embeddings or vision) and an optional role. A partial unique index allows one route per role. brain asks the gateway which route holds a role, so routes can be renamed or rebuilt without touching code, and agents can point at their own.
- Capability is checked twice. The driver must declare it, and the model’s entry in the synced LiteLLM catalog must agree, so an embeddings-only model in a chat chain is skipped before any wire call. Images in the messages add
visionto every candidate. - Sticky beats score. The provider that last served this session successfully on this route goes first, because staying put keeps its prompt cache warm. Sticky is a preference, never an expansion: it must already be in the chain, and editing the route resets it.
- Strategies are per route.
orderedkeeps my written order.auto,priceandlatencyscore entries from declared prices and the last hour of ledger stats. Uptime multiplies the score, so a cheap provider that keeps failing sinks under every strategy. - Rejections are explained. If nothing passes, the
no_routeerror lists every skipped candidate and why.
Challenge 3: when is it safe to switch providers?
If a provider fails, why not just replay the request on the next one?
The rule rests on one contract: every driver translates its wire format into the same channel of events, ending with exactly one terminal event, done or error. Drivers also normalize usage, so input tokens always exclude cache reads.
| Option | How it works | Why not (or why) |
|---|---|---|
| A. No failover | One provider per request | One outage stops every turn. |
| B. Replay anywhere | On any error, resend to the next provider | The client already rendered half an answer, maybe a tool call; a second model writes a different one on top. |
| C. Splice | Continue the partial answer on another provider | Two models writing one answer, and tool-call state cannot be handed over. |
| D. A bright line (chosen) | Fail over only before the first content event | The client sees one provider’s answer, or an honest error. |
I chose the bright line. An attempt’s errors are held back until it produces content: a chunk, reasoning_chunk or tool event. Until then, a failed attempt is drained quietly, recorded, and the loop moves on with a failover event. After that, the stream belongs to that provider, and any error goes to the client as-is.
- Not every error advances. Timeouts, 5xx, 429 and 401/403 (a bad key on this provider) move on. 400, 404, 413 and 422 stop the chain, because the next provider would reject the same request.
- Retries are sized by position. The HTTP drivers retry 429, 5xx and network errors with jittered exponential backoff, but only once while more entries remain; the last entry gets three.
- Exhaustion names every attempt by error code only, never raw provider text.
Challenge 4: how do you account for cost honestly?
Why not store tokens and price them later?
| Option | How it works | Why not (or why) |
|---|---|---|
| A. Price at query time | Store tokens; price them when the dashboard asks | A price edit rewrites history overnight. |
| B. Provider billing exports | Reconcile from each vendor’s usage export | Late, per vendor, blind to sessions and agents. |
| C. Price at request time (chosen) | Cost computed when usage arrives and frozen in the row | The row says what was known then; unknown stays NULL. |
I chose to freeze cost at request time. Every attempt, failed ones included, leaves one cost_ledger row: provider, model, route, agent, purpose, session, tokens, latency, status and cost. The done event carries the ledger_id and cost back to brain. A failed insert is logged, never raised: accounting must not break serving.
- Prices come from two places. An operator’s per-model override wins; otherwise the synced catalog. If neither has the model,
costisNULL, with a warning and an unpriced-usage counter, so the gap is visible instead of silently zero. - Cache writes are split by tier. One-hour cache writes are stored apart from five-minute ones because they bill higher. With no declared one-hour price, the cost is
NULLrather than a guessed multiple. - Currency is stored, never converted. FX rates exist for display only.
- The same rows steer routing. Stickiness, scored strategies, the usage dashboard and daily and monthly budgets all read the ledger.
What broke
Some reasoning-class OpenAI models returned a 200 and an empty stream over chat/completions on turns where a tool call was mandatory. Counting that as success is doubly wrong: the turn looks fine, and stickiness, which reads status='ok' rows, pins the session to the provider that returned nothing. Now a contentless stream is booked empty_output and advances the chain, and a dedicated openai-responses driver serves those models. Success means content, not a status code.
What’s next
Part 3 follows the stream into brain, where every session is an append-only event log with two projections.