System Design September 1, 2026 ~9 min read

Designing a Self-Hosted AI Assistant (Part 1): The Architecture

TL;DR: Timothy is a self-hosted AI assistant I built in Go for one user: me. It gives me one chat interface over every model provider I use, memory in my own database, a ledger of every token spent, and long-running work that continues after I close the tab. The design follows three rules.

First, deterministic code orchestrates, and each model call does exactly one thing. Go decides which model to call, what goes into the prompt, whether a tool may run, and when to stop. A model call answers one question and returns data.

Second, services are split along trust boundaries, not features. The gateway is the path to every model provider, sandboxd alone holds the Docker socket, and brain is the only service with a public API.

Third, one Postgres holds all state. Sessions, memories, vectors, the cost ledger, missions and encrypted secrets share one database with pgvector. No broker, no separate vector database.

This series is written as a design reference: the requirements, the options I compared, and why I chose what I chose. The code is on GitHub.

The problem

What does a personal assistant need that a chat tab does not give?

I use several model providers, each in its own tab with its own history. None shows what a conversation cost across providers, none remembers what I told the others, and none works unless I am typing. Timothy replaces those tabs:

RequirementDetail
UsersExactly one. One bearer token guards everything.
ProvidersAny number, configured at runtime as rows, not code. Hosted APIs and local models side by side.
StateConversations, memory, configuration and costs live in my database.
CostEvery model call lands in a ledger. An unknown price is stored as unknown, never guessed.
Unattended workMissions and automations run without me, and stop for approval before anything destructive.
FootprintOne Docker Compose file on one machine. No Kubernetes, no broker.

Challenge 1: who decides what happens next?

Should the model run the show, or the code?

OptionHow it worksWhy not (or why)
A. Agent frameworkA library supplies the loop, memory, tool calling and prompt templatesPrompts hide behind abstractions. When a turn goes wrong I need the exact bytes sent.
B. Model as orchestratorA planning model picks tools, sub-agents and models, and decides when work is doneIt loses coherence, defends its own mistakes and declares unfinished work done. Every safety rule becomes a prompt it can be talked out of.
C. Go orchestrates, each call does one jobPlain Go calls the gateway. Each model call has one purpose and returns data that code validatesNo framework to lean on; the plumbing is mine to write and test.

Go decisions beside single-purpose model calls
Figure 1: Go makes every decision that must be repeatable, and each model call has one job that the ledger's purpose column records.

I chose C. The deciding argument: a prompt can be persuaded and an if statement cannot. Routing, compaction, memory selection and tool permission must give the same answer every time, so they are Go. Models get the reading, summarising and writing, one narrow job per call.

  • One agent loop, in one place. brain owns the only tool loop; the gateway never loops. The loop has a 16-step ceiling, warns the model one step early, and sends the last step with no tools.
  • Every side call is labelled. Distillation, memory extraction, compaction summaries, titles, agent dispatch and knowledge-base filing are separate calls, each with its own purpose in the cost ledger.

Challenge 2: where do the seams go?

One process, a service per feature, or something else?

OptionHow it worksWhy not (or why)
A. One binaryEverything in one Go processSimplest to run, but the process reading untrusted web pages and model output would also hold every provider key and the Docker socket, which is root on the host.
B. A service per featureChat, missions, automations and the knowledge base as separate services over a brokerMany deploy units for one user, and the seams follow features, which change, not privileges, which do not.
C. Split where privilege or runtime differsA few Go services, each holding one dangerous thing, plus sidecars for Python-only toolsMore containers, and internal calls cross the network.

I chose C. A seam earns its place when it keeps a key, a socket or a runtime away from code that handles untrusted input, and brain handles untrusted input (web pages, email, tool results, model output) all day.

Timothy architecture: every compose service, both networks, Postgres, mission sandboxes and model providers

Figure 2: Every service in the compose file, where only web and brain publish ports and sandboxd shares a network with brain alone. Open full-size diagram

ServiceOwnsNever does
brainpublic API, agent loop, tools, missions, channels, connectorstouch the Docker socket
gatewayproviders, routing, failover, normalized streaming, cost ledgerrun tools or loop
memorydmemories, knowledge base, hybrid searchhold a provider key (it embeds through the gateway)
sandboxdthe Docker socket, one container per missionconnect to the database
sidecarssearxng, markitdown, ocr, pdfgen, whisper (opt-in)touch the database
  • Internal APIs carry no auth, on purpose. They publish no host ports. A shared secret would have to live in brain’s environment, exactly what the split assumes may be compromised, so the network is the boundary.
  • sandboxd sits on its own network, timothy-sandbox, with brain as the only other member. Its API takes a mission ID, never an image, mount or container name, and its container runs read-only with every capability dropped.
  • The gap I accepted. brain also holds the master key, because connectors, channels and delegated coding CLIs need credentials of their own. Every model call from Timothy’s own loop uses a key that exists only in the gateway, but brain is not key-free.

Challenge 3: where does state live?

One database, or the right store for each shape of data?

OptionHow it worksWhy not (or why)
A. A store per shapePostgres for rows, a vector database for embeddings, a broker for eventsThree backups, three failure modes, and a sync job between rows and vectors that will drift.
B. A database per serviceEach service owns its own storeNo shared transactions. brain and the gateway both write the cost ledger, for example.
C. One Postgres with pgvectorEvery table in one database; vectors and full-text columns sit on the same rows as the textOne instance to keep healthy; when it is down, everything is.

I chose C. At one user’s scale Postgres is good enough at every job here, and one database means one backup and transactions that span what would otherwise be separate systems.

  • Vectors live beside their text. memories and kb_chunks each carry a vector(1024) column with an HNSW index and a generated tsvector column, so hybrid search needs no sync job.
  • An inbox table instead of a broker. A side effect that must survive a crash (a finished mission, a cron boundary, a webhook) goes into events in the same transaction as its cause. A drainer delivers it after commit, with bounded retries.
  • Append-only where history matters. session_events, mission_events and memories are never updated in place; a correction is a new row.

How one chat request flows

Sequence of one chat turn across browser, brain, memoryd, gateway and provider
Figure 3: One chat turn, from the browser to a provider and back.

  1. The browser sends POST /v1/sessions/{id}/messages; the web container’s nginx proxies /v1 to brain.
  2. brain claims the session’s single turn slot (a second request gets 409) and appends a user_message event. The turn now belongs to the session, not the HTTP request: closing the tab does not stop it.
  3. brain compacts if the context passed 60% of the model’s window, then projects session_events into model messages.
  4. memoryd embeds the query through the gateway and returns a fenced memory block of at most 1,500 tokens.
  5. brain assembles the system prompt: a stable prefix first, so provider prompt caches keep hitting, and per-turn blocks at the tail.
  6. The loop calls POST /v1/stream on the gateway with a route name. The gateway resolves it to a provider chain, streams, and writes a cost_ledger row per attempt.
  7. Normalized events flow to brain and on to the browser over SSE. Tool calls run in brain through the permission chain, then the loop calls again.
  8. brain appends the assistant_turn. Distillation, memory extraction, a compaction check and a first-turn title run after the answer is durable.

Degrade, never crash

What happens when a dependency is not there?

One machine restarts in odd orders: Postgres comes up late, a sidecar is off, a provider is down. Every Go service serves /health before its database is reachable. The pool reconnects with backoff, migrations retry, and /health reports degraded with a reason instead of the container restart-looping.

Optional features follow the same rule: without the OCR sidecar a knowledge-base image gets no description rather than a failed ingest, and an unset WHISPER_URL leaves transcription unmounted. Only configuration that cannot fix itself fails fast: a missing master key stops the gateway at boot.

What’s next

Part 2 opens the gateway: providers as data, routes as roles, one event stream for every provider, and the bright line that decides when failover is still safe.

// end of article — process exited with code 0

// share:Twitter / XLinkedIn
// llm:View as Markdown

// comments

loading...